A practical benchmark evaluation of four July 2026 LLM releases — Kimi K3, Claude Opus 5, Grok 4.5, and Gemini 3.6 Flash — using Baba Is Bench, an agent benchmark based on the puzzle game Baba Is You. Results show Claude Opus 5 performs nearly as well as the more expensive Claude Fable 5 at lower cost. Kimi K3 is competitive and now open-weight on Hugging Face. Grok 4.5 underperforms relative to official benchmarks, suggesting possible overfitting. Gemini 3.6 Flash is a disappointment: it loops endlessly on unsolvable levels, burning over $260 in API costs and failing to improve on its predecessor in either skill or pricing. The post also highlights tooling challenges, particularly around finding a neutral harness for fair cross-model evaluation.

6m read timeFrom quesma.com
Post cover image
Table of contents
Stage 0: The IntroStage 1: The LakeConclusion
400 Impressions