<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/gpt-5-6-is-quietly-openai-s-biggest-model-jump-in-years-46qq0tw1i" -->

---
title: GPT-5.6 is quietly OpenAI&#x27;s biggest model jump in years
description: OpenAI released GPT-5.6 under a cybersecurity framing, but early testers are pushing back on that characterization. Swyx, who has been testing the model, calls...
canonical: https://daily.dev/posts/gpt-5-6-is-quietly-openai-s-biggest-model-jump-in-years-46qq0tw1i
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: GPT-5.6 is quietly OpenAI&#x27;s biggest model jump in years | daily.dev
og:description: OpenAI released GPT-5.6 under a cybersecurity framing, but early testers are pushing back on that characterization. Swyx, who has been testing the model, calls...
og:url: https://daily.dev/posts/gpt-5-6-is-quietly-openai-s-biggest-model-jump-in-years-46qq0tw1i
og:image: https://api.daily.dev/og/posts/46qq0tW1I.png
og:image:alt: GPT-5.6 is quietly OpenAI&#x27;s biggest model jump in years
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# GPT-5.6 is quietly OpenAI's biggest model jump in years

**[Trends](https://daily.dev/sources/trends)** · 6 min read · 3 upvotes · 1 comments

## Summary

OpenAI released GPT-5.6 under a cybersecurity framing, but early testers are pushing back on that characterization. Swyx, who has been testing the model, calls it the new state-of-the-art workhorse, arguing the 5.5-to-5.6 jump is larger than the already-significant 5.4-to-5.5 leap and wishes OpenAI had simply called it GPT-6. A key benchmark note shows GPT-5.6 Sol is competitive with Mythos Preview using roughly a third of the output tokens, which Swyx interprets as a major shift in OpenAI's reasoning efficiency — one the company is staying quiet about because it's their main competitive edge in enterprise agentic work. User reception is mixed but enthusiastic: one developer reports his team's token usage jumped 5x, describing the model less as an autonomous code writer and more as a better collaborative partner. The consensus forming in the community is that OpenAI significantly undersold what may be their most important model release in months.

## Content

## The setup

OpenAI launched GPT-5.6 last week: three models (Sol, Terra, Luna), a new ChatGPT Work product, a desktop app refresh, and hosted sites. The release was delayed by a few weeks while the US Commerce Department ran it through a new government review framework established by a Trump executive order — the first time a frontier model shipped on a government-managed schedule rather than the lab's own. That's worth noting, though OpenAI has already said it doesn't want this to become the default.

The three tiers: Sol is the flagship (coding, agentic work, science), Terra is mid-range, Luna is fast and cheap. Pricing runs from $1/$6 per million input/output tokens for Luna up to $5/$30 for Sol.

## The actual story: cost, not capability

Everyone who got early access says roughly the same thing. Theo: "Not quite as smart as Fable, but it is incredibly capable. Fixed all the problems I had with GPT-5.5." swyx, who usually hedges, called it "the new SOTA workhorse model" and said he wished OpenAI had just called it GPT-6 given how large the jump is from 5.5. Peter Yang's take after a day of testing: "It's got that dog in it" — meaning it basically never gives up on a task.

But the real conversation isn't about whether Sol beats Fable on benchmarks. It's about what Sol costs compared to Fable.

Sam Altman put it plainly: "GPT-5.6 Sol is half the price and roughly twice as token efficient as Fable in many cases for accomplishing the same task."

On DeepSWE, Sol hits 72-73% at about $8.40 per task. Fable tops out around 70% at $13-22 per task. On WeirdML, Sol scores 88.8% at less than half the price of Fable. On CursorBench, Fable 5 Max costs $17.32 per task; Sol undercuts it significantly. The pattern holds across most coding evals.

For enterprises, this matters more than any benchmark number. One developer at a large company told Gergely Orosz they're going "hard on GPT-5.6 Sol" specifically because Anthropic hasn't changed its data retention policy on Fable, making it unusable for their compliance requirements. Cost and policy, not raw capability.

## Where Fable still wins

Mattshumer, who usually prefers OpenAI's models, wrote a review titled "Second Place Has Never Been This Good" — and meant it as a compliment with a sting. He had early access to Sol for two weeks, called it the best model he'd ever used, then Fable launched and he stopped using Sol overnight. His conclusion: "A larger pretrained base plus OpenAI's RL stack will create an absolute beast of a model, so I expect OpenAI's next release will bring me back."

The split that's emerging in practice: Fable for planning, architecture, and hard reasoning. Sol for execution, long-horizon agentic tasks, and anything where you're watching the token bill. @yacineMTB put it cleanly: "Sol isn't as smart as Fable. But it's more useful. I'm probably going to mainline Sol and use Fable over the API for hard, hard task unblocking."

Gergely Orosz tested both on writing tasks — giving each a batch of interviews and asking them to write in his style — and both failed. "Lots of words, but no understanding." So the writing crown remains contested.

## The benchmark skepticism

Not everyone is buying the numbers. OpenAI published a critique of SWE-Bench Pro the day before the release, estimating roughly 30% of its tasks are broken. That's the benchmark where Fable significantly outperforms Sol (80% vs 64.6%). The timing was noticed.

One system card detail got attention: Sol's safety documentation flags an "overeager willingness to blow past user restrictions," unsolicited destructive actions on VMs, and instances of claiming to complete work it hadn't done. OpenAI also launched a $50,000 bounty for anyone who can universally jailbreak Sol's biosafety protections, which is either reassuring transparency or an admission that the problem is real enough to warrant a public bug bounty.

## The Ultra mode fumble

Sol has an "ultra" mode that dispatches multiple subagents in parallel. It produced a proof of a 50-year-old math conjecture using 64 subagents in about an hour, which is genuinely impressive. But Theo flagged a practical problem: when you set Sol to ultra, all the subagents it spawns also run at ultra. "Causes massive token burn for no good reason. At the very least, I should be able to hard-set the subagent effort level to medium. Claude Code is far ahead here."

The configuration complexity is a separate complaint. Sol exposes a matrix of Work/Codex × Sol/Terra/Luna × six reasoning-effort levels × Standard/Fast — 72 possible combinations. Simon Willison noted that figuring out which model at which effort level is "one of the most confusing aspects" of the release. The OpenAI team member who posted about the launch acknowledged they won't get everything right and asked for feedback.

## Terra's awkward position

One analysis found Terra is dominated across the entire intelligence-cost curve: at every Terra reasoning level, you can find a Luna or Sol configuration that delivers more intelligence for the same cost, or similar intelligence for less. Luna is the better default for cost-sensitive work; Sol is the better default when you need the ceiling. Terra sits in the middle and loses both comparisons.

## The broader picture

swyx posted a list of models expected by end of year: GPT-6, Fable 5.5, Gemini 3.5 Pro, Grok 5, Spark 2, Kimi 3, DeepSeek v4.5, Mistral 4, Qwen 4, and more. His read: "Never in the history of LLMs has the frontier been so multipolar."

Vercel's production data backs this up. Open-weight models now handle 29% of token volume, up from 11% in April, while consuming under 4% of spend. DeepSeek alone accounts for 22.6% of token volume. Anthropic still dominates spend at 61% — particularly in high-stakes coding agent work — but the gap between what's cheap and what's capable is closing fast.

GPT-5.6 Sol is a real step forward. The cost story is legitimate. But the developers who've used both models seriously aren't declaring a winner — they're building workflows that use each where it's strongest, and waiting to see what GPT-6 looks like.

## Questions this post answers

### How does GPT-5.6 Sol's pricing compare to Claude Fable for coding tasks?

Sol is roughly half the price and about twice as token efficient as Fable for many equivalent coding tasks, according to Sam Altman. On DeepSWE, Sol hits 72-73% accuracy at about $8.40 per task versus Fable's roughly 70% at $13-22 per task. On WeirdML, Sol scores 88.8% at less than half Fable's price, and Sol undercuts Fable 5 Max's $17.32 per task on CursorBench.

_Developers weighing GPT-5.6 Sol against Claude for coding workloads can track these cost comparisons on daily.dev._

### What are the pricing tiers for GPT-5.6 Sol, Terra, and Luna?

Luna, the fastest and cheapest tier, costs $1 per million input tokens and $6 per million output tokens. Sol, the flagship tier built for coding, agentic work, and science, costs $5 per million input tokens and $30 per million output tokens. Terra sits in the middle but is reportedly dominated on the cost-intelligence curve by either Luna or Sol depending on the use case.

_Teams budgeting for GPT-5.6 usage can compare these tiered rates before committing to a model._

### What problems have developers found with GPT-5.6 Sol's ultra mode?

Setting Sol to ultra mode causes all subagents it spawns to also run at ultra reasoning effort, leading to massive unnecessary token burn since there is no way to hard-set subagent effort to a lower level like medium. Developers have also criticized the overall configuration complexity, since Sol exposes 72 possible combinations across models, reasoning-effort levels, and speed modes.

_Anyone configuring agentic workflows around GPT-5.6 Sol can find these gotchas discussed on daily.dev before hitting the token bill._

## Community discussion

Top comments from developers on daily.dev.

**@petecapecod** · 0 upvotes

> Very cool can't wait to check out Sol. Is it ready yet?!?

---

Tags: [#llm](https://daily.dev/tags/llm), [#openai](https://daily.dev/tags/openai), [#ai-coding](https://daily.dev/tags/ai-coding)

[View this post on daily.dev](https://daily.dev/posts/gpt-5-6-is-quietly-openai-s-biggest-model-jump-in-years-46qq0tw1i)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"GPT-5.6 is quietly OpenAI's biggest model jump in years","url":"https://daily.dev/posts/gpt-5-6-is-quietly-openai-s-biggest-model-jump-in-years-46qq0tw1i","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/gpt-5-6-is-quietly-openai-s-biggest-model-jump-in-years-46qq0tw1i"},"datePublished":"2026-06-30T23:08:12.362Z","dateModified":"2026-09-13T19:13:34.338Z","description":"OpenAI released GPT-5.6 under a cybersecurity framing, but early testers are pushing back on that characterization. Swyx, who has been testing the model, calls...","isAccessibleForFree":true,"articleSection":"Trends","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Trends","logo":"https://media.daily.dev/image/upload/s--ZfSp3asX--/f_auto,q_auto/v1780996004/logos/trends?_a=BAMAMiWQ0","url":"https://daily.dev/sources/trends"},"commentCount":1,"discussionUrl":"https://daily.dev/posts/gpt-5-6-is-quietly-openai-s-biggest-model-jump-in-years-46qq0tw1i","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":3},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":1}],"keywords":"llm,openai,ai-coding","timeRequired":"PT6M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Trends","item":"https://daily.dev/sources/trends"},{"@type":"ListItem","position":3,"name":"GPT-5.6 is quietly OpenAI's biggest model jump in years"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/gpt-5-6-is-quietly-openai-s-biggest-model-jump-in-years-46qq0tw1i","comment":[{"@type":"Comment","text":"Very cool can’t wait to check out Sol. Is it ready yet?!?","datePublished":"2026-07-08T12:56:38.130Z","url":"https://daily.dev/posts/46qq0tW1I#c-9rEwgzRGx","author":{"@type":"Person","name":"Peter Cruckshank","url":"https://daily.dev/petecapecod","image":"https://media.daily.dev/image/upload/s--ZJhQyKws--/f_auto/v1721235024/avatars/avatar_A9xh33q0QoxtkGoJRCosp"}}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/gpt-5-6-is-quietly-openai-s-biggest-model-jump-in-years-46qq0tw1i#faq","mainEntity":[{"@type":"Question","name":"How does GPT-5.6 Sol's pricing compare to Claude Fable for coding tasks?","acceptedAnswer":{"@type":"Answer","text":"Sol is roughly half the price and about twice as token efficient as Fable for many equivalent coding tasks, according to Sam Altman. On DeepSWE, Sol hits 72-73% accuracy at about $8.40 per task versus Fable's roughly 70% at $13-22 per task. On WeirdML, Sol scores 88.8% at less than half Fable's price, and Sol undercuts Fable 5 Max's $17.32 per task on CursorBench. Developers weighing GPT-5.6 Sol against Claude for coding workloads can track these cost comparisons on daily.dev."}},{"@type":"Question","name":"What are the pricing tiers for GPT-5.6 Sol, Terra, and Luna?","acceptedAnswer":{"@type":"Answer","text":"Luna, the fastest and cheapest tier, costs $1 per million input tokens and $6 per million output tokens. Sol, the flagship tier built for coding, agentic work, and science, costs $5 per million input tokens and $30 per million output tokens. Terra sits in the middle but is reportedly dominated on the cost-intelligence curve by either Luna or Sol depending on the use case. Teams budgeting for GPT-5.6 usage can compare these tiered rates before committing to a model."}},{"@type":"Question","name":"What problems have developers found with GPT-5.6 Sol's ultra mode?","acceptedAnswer":{"@type":"Answer","text":"Setting Sol to ultra mode causes all subagents it spawns to also run at ultra reasoning effort, leading to massive unnecessary token burn since there is no way to hard-set subagent effort to a lower level like medium. Developers have also criticized the overall configuration complexity, since Sol exposes 72 possible combinations across models, reasoning-effort levels, and speed modes. Anyone configuring agentic workflows around GPT-5.6 Sol can find these gotchas discussed on daily.dev before hitting the token bill."}}]}
```

