GitHub shares benchmark results comparing the GitHub Copilot agentic harness against model-vendor harnesses (Claude Code and Codex CLI) across five benchmarks: SWE-bench Verified, SWE-bench Pro, SkillsBench, TerminalBench, and an internal Win-Hill benchmark. Using four models (Claude Sonnet 4.6, Claude Opus 4.7, GPT-5.4, GPT-5.5), the Copilot harness achieves task resolution rates on par with vendor harnesses while consuming fewer tokens across most configurations. A key differentiator is multi-model flexibility — supporting 20+ frontier models — enabling users to trade off cost vs. peak quality per task. The post also details methodology, including five independent runs per configuration and controlled normalization of context windows and reasoning effort.

8m read timeFrom github.blog
Post cover image
Table of contents
More optimizations we are makingHow we iterate with benchmarksToken efficiencyTask resolutionTerminalBench: Token efficiency, task completion, and varianceOne harness, many modelsConclusionTry it yourselfMethodologyTags:Written by
473 Impressions