AI agents pass tests. Can they write idiomatic Laravel?
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
AI coding agents using Laravel Boost can now pass all 17 Laravel eval tasks at or near 100% accuracy, effectively saturating the existing benchmark. With correctness becoming table stakes, the focus is shifting to two new metrics: correct code per token (efficiency) and idiomatic Laravel quality. The post outlines how Boost will evolve to measure whether agents write code that not only works but follows Laravel conventions — using form requests, eager loading, Route::resource(), and other best practices — while minimizing the context tokens needed to get there. A new scoring layer using LLM-as-judge and a 19-point best-practices rubric is being developed alongside token/cost reporting as headline metrics.