AI benchmarks are systematically broken by two forces: training data contamination (models memorize test questions from public datasets) and adversarial gaming (labs submit best-of-N variants to leaderboards). GSM1k showed accuracy drops of up to 13% when models faced unseen math problems, exposing memorization rather than reasoning. MMLU contains ~6.5% erroneous questions, and Chatbot Arena was gamed by Meta submitting 27 private Llama-4 variants. The root cause is Goodhart's Law: once a benchmark score drives funding and procurement, it stops measuring what it was designed to measure. Practical mitigations include private rotating test sets, contamination detection, and building small domain-specific evaluations from your own data rather than trusting public leaderboards.

10m read timeFrom cacm.acm.org
Post cover image
Table of contents
An Old Problem in New HardwareContamination: The Cheapest Way to CheatThe Arena: When the Judge Gets Gamed, TooWhy This Keeps HappeningWhat Actually Survives GoodhartThe Score is Partly Fiction. Decide accordingly.
2.5K Impressions1 Comment