Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
A benchmark study testing frontier LLMs (Claude Fable 5, GPT-5.6 Sol, Gemini 3.5 Flash, GLM-5.2, and others) on the puzzle game Baba Is You, ported to the Harbor agent evaluation framework. Claude Fable 5 and GPT-5.6 Sol solved 13 of 14 levels in Stage 1, but both were 4x and 13x slower respectively than a human Twitch streamer playing for the first time. Key findings: token price is misleading — total cost per task matters more; Gemini 3.5 Flash was 2.4x more expensive than Fable 5 despite lower token prices due to excessive turns; GPT-5.6 Sol was the most cost-efficient frontier model; open-weight GLM-5.2 outperformed both Gemini models. The team spent over $2,000 running experiments.