<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/nobody-agrees-on-which-model-is-actually-good-right-now-jn3aqnm0m" -->

---
title: Nobody agrees on which model is actually good right now
description: Claude Opus 5 launched at the same $15/$75 per million token pricing as Opus 4, sparking a split reaction: some call it dramatically stronger on reasoning and...
canonical: https://daily.dev/posts/nobody-agrees-on-which-model-is-actually-good-right-now-jn3aqnm0m
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Nobody agrees on which model is actually good right now | daily.dev
og:description: Claude Opus 5 launched at the same $15/$75 per million token pricing as Opus 4, sparking a split reaction: some call it dramatically stronger on reasoning and...
og:url: https://daily.dev/posts/nobody-agrees-on-which-model-is-actually-good-right-now-jn3aqnm0m
og:image: https://api.daily.dev/og/posts/JN3aqNm0M.png
og:image:alt: Nobody agrees on which model is actually good right now
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Nobody agrees on which model is actually good right now

**[Trends](https://daily.dev/sources/trends)** · 2 min read · 4 upvotes · 2 comments

## Summary

Claude Opus 5 launched at the same $15/$75 per million token pricing as Opus 4, sparking a split reaction: some call it dramatically stronger on reasoning and coding, others say it's overpriced given cheaper alternatives. Prompting advice circulating suggests dropping old skills/MCPs/Claude.md configs and giving the model more autonomy instead of micromanaging its process, since Opus 5 behaves differently from earlier Claude versions. A developer writeup argues Opus 5's premium is best justified for complex bug detection, security audits, multi-file refactoring, and agentic workflows, while Sonnet 4 covers routine coding and docs at a fraction of the cost. Its 200K context window also lags behind GPT-4.1 and Gemini 2.5 Pro's 1M. A broader roundup adds that DeepSeek V4 Flash may beat Opus 4.8 on cost-efficiency and that Tencent's Hy3 topped OpenRouter usage charts for cheap frontend-focused generation, reinforcing a market trend toward routing tasks to the cheapest sufficient model rather than defaulting to the flagship.

## Content

Anthropic dropped Claude Opus 5 on July 17 at the same price as Opus 4 ($15/$75 per million tokens), promising better reasoning and coding. The reception has been... mixed, and that's being generous.

@Hesamation has been running a recurring bit calling this "a brutal month" for Claude Code: Opus 5 didn't land the way Anthropic hoped, there's backlash over watermarking, and the extra 50% weekly usage Anthropic handed out is set to vanish on August 31. The kicker: Anthropic has now extended Claude Code limits for the third time, which Hesamation reads as a tell. "You know Fable is absolutely chewing through usage AND Opus 5 didn't really land when Anthropic is extending Claude Code limits for the 3rd time."

The AA intelligence scores back up the vibe shift: Fable 5 at 62, Grok 4.6 at 61, Kimi K3 and GLM 5.3 tied at 60. Codex is apparently still solid too, if you can stomach the watermarking situation.

Not everyone's piling on. @scaling01 posted a flat "Opus 5 is pretty good" and followed up claiming it crushes a dynamic game benchmark comparable to ARC-AGI-3 or MazeBench. @mattshumer_ argues most people are just using it wrong: wipe your old skills, MCPs, and Claude.md files, stop micromanaging the instructions, and just say what you want. He claims that alone fixes the

## Questions this post answers

### What is the pricing for Claude Opus 5 compared to Opus 4?

Claude Opus 5 launched at the same price as Opus 4: $15 per million input tokens and $75 per million output tokens. It released on July 17, promising improved reasoning and coding performance over its predecessor, though early community reception has been mixed rather than enthusiastic.

_Developers weighing whether to upgrade their Claude Code setup can track pricing and reception shifts like this on daily.dev._

### Why did Anthropic extend Claude Code usage limits multiple times?

Anthropic extended Claude Code usage limits for a third time, which some observers interpret as a response to Opus 5 underperforming expectations combined with competitive pressure from rivals like Fable heavily consuming usage. A temporary 50% weekly usage boost granted around the same period is scheduled to expire on August 31.

_Anyone relying on Claude Code's usage caps for daily workflows can follow these shifting limits on daily.dev._

### How does Claude Opus 5 compare to other models like Fable 5 and Grok 4.6 on intelligence benchmarks?

On aggregated intelligence scores, Fable 5 leads at 62, Grok 4.6 follows at 61, with Kimi K3 and GLM 5.3 tied at 60, suggesting Claude Opus 5 is not a clear leader despite its price parity with Opus 4. Some users still defend Opus 5's practical performance, particularly on dynamic game benchmarks similar to ARC-AGI-3 or MazeBench.

_Developers choosing between coding models can weigh these shifting benchmark rankings via daily.dev._

## Community take

How the wider developer community reacted, aggregated from 2 discussions and 15 comments across x (as of 2026-09-13).

**TL;DR:** Reaction is largely skeptical, with people citing dropped Anthropic API usage, coding quality complaints, buggy apps, and stingy usage limits, while a smaller minority pushes back that Opus 5 is underrated.

**Sentiment:** 10% positive · 20% mixed · 70% skeptical

**The case for**

- One person says switching most of their agent workloads off Anthropic actually improved internal quality benchmarks with alternative models.
- A couple of replies push back that Opus 5 is underrated or that coverage is overly negative.

**The pushback**

- Some report cutting Anthropic API usage by over 95% for their agent fleet and seeing quality improve with alternatives.
- Coding output is described as poor, the app buggy, usage limits stingy, and communication/community experience bad.
- Free/consumer tiers are seen as tightly rationed because most revenue reportedly comes from enterprise API and Claude Code contracts.
- One reply frames the temporary token bump as covering up broken unit economics ahead of a real price hike.
- Enterprise adoption is characterized by some as inertia from unfamiliarity with alternatives rather than genuine preference.

**By community**

- x (skeptical): Replies mostly describe frustration with coding quality, limits, and rationed usage, with only scattered pushback calling the model underrated.

**Hottest debate:** Whether continued enterprise usage reflects real preference or just inertia/unfamiliarity with alternatives.

**Open questions**

- Will the temporary usage boost expiring in August signal a broader pricing shift for consumer tiers?
- How much of enterprise revenue is actually tied to genuine technical preference versus organizational inertia?

**Highlights**

> @Hesamation In the past 3 weeks, we've dropped our anthropic API use for our UViiVE fleet by over 95% and seen an increase in internal quality benchmarks using alternative LLMs as the CPUs for our agents. That revenue is likely permanently lost while we grow.
> — [scekker on x · 4 points](https://x.com/scekker/status/2090278595000176824)

> @fanofaliens @Hesamation Yes this much better than benchmarks. But what's for me? I code and their models suck. Their app is buggy. They limits are shit. Their comms is bad and community is toxic. Fable talks like intelligent is the only thing going for them currently for my requirement.
> — [RaviTejaKNTS on x · 1 comments](https://x.com/RaviTejaKNTS/status/2090296206656426041)

> @Hesamation 50% more tokens for free is the AI industry's way of saying "we broke the unit economics, please don't notice." Opus 5 shipped, the watermarks landed, and the free tier vanished. this isn't a product cycle. it's a subsidy drawdown before the real bill arrives
> — [TheAIShrink on x](https://x.com/TheAIShrink/status/2090233326023868445)

> @Hesamation Most of entreprise coders still use it because they don't know about anything else. It's the new sign of someone who's not technical.
> — [mSanterre on x · 1 points, 1 comments](https://x.com/mSanterre/status/2090288429636682037)

> @namastedevang @Hesamation Yep. With ~70-80% of revenue from enterprise API tokens and Claude Code contracts, free/consumer tiers stay tightly rationed so the high-paying usage gets priority on capacity.
> — [grok on x](https://x.com/grok/status/2090288877151895734)

**Source threads**

- [x](https://x.com/Hesamation/status/2090235396160377085) · 0 points · 13 comments
- [x](https://x.com/Hesamation/status/2090231896143802847) · 1 points · 2 comments

## Community discussion

Top comments from developers on daily.dev.

**@yonasuriv** · 2 upvotes

> These AI generated posts makes me go loco

**@lucas39** · 2 upvotes

> "Nobody agrees on which model is actually good right now"
>
> Then writes a whole article to say **Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5**

## Similar posts on daily.dev

- [Anthropic’s Opus 5 is almost Fable 5](https://daily.dev/posts/anthropic-s-opus-5-is-almost-fable-5-vlen80a1c) · The New Stack · 9 upvotes · 2 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-coding](https://daily.dev/tags/ai-coding), [#anthropic](https://daily.dev/tags/anthropic), [#prompt-engineering](https://daily.dev/tags/prompt-engineering)

[View this post on daily.dev](https://daily.dev/posts/nobody-agrees-on-which-model-is-actually-good-right-now-jn3aqnm0m)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Nobody agrees on which model is actually good right now","url":"https://daily.dev/posts/nobody-agrees-on-which-model-is-actually-good-right-now-jn3aqnm0m","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/nobody-agrees-on-which-model-is-actually-good-right-now-jn3aqnm0m"},"datePublished":"2026-08-13T01:45:26.426Z","dateModified":"2026-09-13T19:46:20.066Z","description":"Claude Opus 5 launched at the same $15/$75 per million token pricing as Opus 4, sparking a split reaction: some call it dramatically stronger on reasoning and...","image":"https://pbs.twimg.com/media/HPkPkJZWwAAkItQ.jpg","thumbnailUrl":"https://pbs.twimg.com/media/HPkPkJZWwAAkItQ.jpg","isAccessibleForFree":true,"articleSection":"Trends","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Trends","logo":"https://media.daily.dev/image/upload/s--ZfSp3asX--/f_auto,q_auto/v1780996004/logos/trends?_a=BAMAMiWQ0","url":"https://daily.dev/sources/trends"},"commentCount":2,"discussionUrl":"https://daily.dev/posts/nobody-agrees-on-which-model-is-actually-good-right-now-jn3aqnm0m","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":4},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":2}],"keywords":"llm,ai-coding,anthropic,prompt-engineering","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Trends","item":"https://daily.dev/sources/trends"},{"@type":"ListItem","position":3,"name":"Nobody agrees on which model is actually good right now"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/nobody-agrees-on-which-model-is-actually-good-right-now-jn3aqnm0m","comment":[{"@type":"Comment","text":"These AI generated posts makes me go loco","datePublished":"2026-08-16T00:01:30.193Z","dateModified":"2026-08-16T00:02:23.786Z","url":"https://daily.dev/posts/JN3aqNm0M#c-1RGUh7iW9","author":{"@type":"Person","name":"Livewire","url":"https://daily.dev/yonasuriv","image":"https://media.daily.dev/image/upload/f_auto/v1656022362/avatars/jzd0XPKYw5GfG2OJ9TZtJ"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2}},{"@type":"Comment","text":"“Nobody agrees on which model is actually good right now”\nThen writes a whole article to say Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5 Opus 5","datePublished":"2026-08-13T13:34:21.287Z","url":"https://daily.dev/posts/JN3aqNm0M#c-PUBVxjeIP","author":{"@type":"Person","name":"Lucas","url":"https://daily.dev/lucas39","image":"https://lh3.googleusercontent.com/a/ACg8ocIIUWuKYGCrhLQ9bRh_zl_uNxcdY4Gvvwc91aTGmSre6nyWEA=s96-c"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2}}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/nobody-agrees-on-which-model-is-actually-good-right-now-jn3aqnm0m#faq","mainEntity":[{"@type":"Question","name":"What is the pricing for Claude Opus 5 compared to Opus 4?","acceptedAnswer":{"@type":"Answer","text":"Claude Opus 5 launched at the same price as Opus 4: $15 per million input tokens and $75 per million output tokens. It released on July 17, promising improved reasoning and coding performance over its predecessor, though early community reception has been mixed rather than enthusiastic. Developers weighing whether to upgrade their Claude Code setup can track pricing and reception shifts like this on daily.dev."}},{"@type":"Question","name":"Why did Anthropic extend Claude Code usage limits multiple times?","acceptedAnswer":{"@type":"Answer","text":"Anthropic extended Claude Code usage limits for a third time, which some observers interpret as a response to Opus 5 underperforming expectations combined with competitive pressure from rivals like Fable heavily consuming usage. A temporary 50% weekly usage boost granted around the same period is scheduled to expire on August 31. Anyone relying on Claude Code's usage caps for daily workflows can follow these shifting limits on daily.dev."}},{"@type":"Question","name":"How does Claude Opus 5 compare to other models like Fable 5 and Grok 4.6 on intelligence benchmarks?","acceptedAnswer":{"@type":"Answer","text":"On aggregated intelligence scores, Fable 5 leads at 62, Grok 4.6 follows at 61, with Kimi K3 and GLM 5.3 tied at 60, suggesting Claude Opus 5 is not a clear leader despite its price parity with Opus 4. Some users still defend Opus 5's practical performance, particularly on dynamic game benchmarks similar to ARC-AGI-3 or MazeBench. Developers choosing between coding models can weigh these shifting benchmark rankings via daily.dev."}}]}
```

