A head-to-head evaluation of Kimi K3 and GPT-5.6 Sol on the DeepSWE software engineering benchmark across 904 rollouts. GPT-5.6 Sol leads on pass@1 (72.7% vs 68.5%) and reliability, but Kimi K3 wins pass@4 (89.4% vs 85.8%) at 2.8x more solved tasks per dollar. The models fail in distinct ways — Sol breaks existing tests more often, Kimi K3 produces near-misses — and have only 0.46 correlation on which tasks they solve. A Kimi-first cascade that escalates to Sol when tests fail reaches ~85.6%, beating either model alone. Together they cover 108 of 113 tasks (95.6%). The post also promotes running Kimi K3 on Together AI's inference platform.

6m read timeFrom together.ai
Post cover image
Table of contents
The DeepSWE scoreboard: pass@1 and pass@kCost comparison: Kimi K3 vs GPT-5.6 Sol pricingKimi K3 vs GPT-5.6 Sol by programming languageCoverage vs reliability: opposite corners of the planeHow different are Kimi K3 and GPT-5.6 Sol?How high can routing between them get you?Where each wins, by task typeFailure modesRun Kimi K3 on Together AIKimi K3 vs GPT-5.6 Sol FAQ
600 Impressions