AI code benchmarks lied to us
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
A deep critique of existing AI coding benchmarks, particularly SWE-bench Pro, arguing they are contaminated, poorly designed, and produce misleading results. The author introduces DeepSWE bench (by Data Curve, a company they've invested in) as a more realistic alternative that tests models on novel tasks across multiple languages with handwritten verifiers. Key findings: GPT-4.5/o3 (called GPT55/54) dramatically outperform other models, Gemini Flash performs far worse than SWE-bench Pro suggests, and Claude Opus uses 2-3x more tokens at higher cost for lower scores. The benchmark reveals a massive gap between open-weight models and frontier models that older benchmarks obscured. The author also encourages developers to build their own mini-benchmarks from real failure cases.