Oh no (the new Grok model is good)

This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).

Grok 4.5 from xAI (now partnered with Cursor) has launched with strong benchmark performance, placing third on DeepSWE behind Claude and GPT-5.5, while being significantly cheaper at $2/M input and $6/M output tokens. The model shows impressive token efficiency — using only 2M tokens per coding task versus 7-9M for competitors — and performs well on multi-step agentic coding tasks. Real-world testing showed it handles complex back-and-forth PR workflows, code auditing, and even 3D game generation better than expected. However, it falls short of the newest model generation (Claude and GPT-5.6) in orchestrating sub-agents. A notable transparency moment: Cursor's codebase was accidentally included in training data, tainting the Cursor Bench results. Overall, xAI has made a remarkable leap from being largely irrelevant to a genuine competitor in the frontier model space.

24m watch time
5 Impressions