GitHub shares benchmark results comparing the GitHub Copilot agentic harness against model-vendor harnesses (Claude Code and Codex CLI) across five benchmarks: SWE-bench Verified, SWE-bench Pro, SkillsBench, TerminalBench, and an internal Win-Hill benchmark. Using four models (Claude Sonnet 4.6, Claude Opus 4.7, GPT-5.4, GPT-5.5), the Copilot harness achieves task resolution rates on par with vendor harnesses while consuming fewer tokens across most configurations. A key differentiator is multi-model flexibility — supporting 20+ frontier models — enabling users to trade off cost vs. peak quality per task. The post also details methodology, including five independent runs per configuration and controlled normalization of context windows and reasoning effort.