<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/openai-s-new-ultrafast-mode-runs-gpt-5-6-sol-14x-faster-using-cerebras-chips-kewuoa7pc" -->

---
title: OpenAI&#x27;s new Ultrafast mode runs GPT-5.6 Sol 14x faster...
description: OpenAI is previewing a new API tier called Ultrafast that runs GPT-5.6 Sol up to 14x faster than standard inference, reaching up to 750 output tokens per...
canonical: https://daily.dev/posts/openai-s-new-ultrafast-mode-runs-gpt-5-6-sol-14x-faster-using-cerebras-chips-kewuoa7pc
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: OpenAI&#x27;s new Ultrafast mode runs GPT-5.6 Sol 14x faster using Cerebras chips | daily.dev
og:description: OpenAI is previewing a new API tier called Ultrafast that runs GPT-5.6 Sol up to 14x faster than standard inference, reaching up to 750 output tokens per...
og:url: https://daily.dev/posts/openai-s-new-ultrafast-mode-runs-gpt-5-6-sol-14x-faster-using-cerebras-chips-kewuoa7pc
og:image: https://api.daily.dev/og/posts/KEwuoA7Pc.png
og:image:alt: OpenAI&#x27;s new Ultrafast mode runs GPT-5.6 Sol 14x faster using Cerebras chips
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI's new Ultrafast mode runs GPT-5.6 Sol 14x faster using Cerebras chips

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 0 upvotes · 0 comments

## Summary

OpenAI is previewing a new API tier called Ultrafast that runs GPT-5.6 Sol up to 14x faster than standard inference, reaching up to 750 output tokens per second. The speedup comes from a partnership with Cerebras, whose Wafer-Scale Engine chips keep 44 GB of SRAM on-chip to avoid memory bandwidth bottlenecks. Cerebras claims no quality tradeoff: GPT-5.6 Sol finished Humanity's Last Exam in 11h11m versus Claude Fable 5's 78h27m, and shows a 5.6x speedup on GDP-Val. The tier targets latency-sensitive enterprise use cases like incident response and customer support, and is rolling out to a small group of customers first, with scaling to broader availability still an open question given the difficulty of scaling Cerebras hardware.

## Content

OpenAI is previewing a new API tier called Ultrafast that runs GPT-5.6 Sol up to 14 times faster than standard inference — around 750 output tokens per second. The trick isn't a smaller or dumbed-down model. It's a change of hardware: instead of racks of Nvidia GPUs, Ultrafast runs on Cerebras's wafer-scale chips, which keep 44 GB of SRAM directly on-chip. That sidesteps the memory bandwidth bottleneck that normally caps how fast you can pull tokens out of a large model.

The preview opened August 13 to a small set of customers, including Jane Street, Podium, Basis, and Rogo. OpenAI says access will widen "as capacity grows," which is the usual way of saying: don't expect this everywhere at once. Pricing hasn't been published yet either.

## Why speed matters here

750 tokens/sec (and reportedly up to ~1,300 tokens/sec on Cerebras's CS-4 hardware, according to one report) isn't just a number to brag about. It changes what's practical to build. Incident response, fraud detection, real-time voice support, and agentic workflows all get bottlenecked by latency — an agent that takes 10 seconds to think isn't usable in a live phone call or a fraud check that needs to happen before a transaction clears. Cut that to under a second, and suddenly things that were demos become products.//

Cerebras also published benchmark numbers worth sitting with: GPT-5.6 Sol Ultrafast reportedly finished all 2,500 questions in Humanity's Last Exam in 11 hours 11 minutes, versus 78 hours 27 minutes for Claude Fable 5. It also showed a 5.6x speedup on GDP-Val with, they claim, no quality loss. I'd want to see independent verification before taking the

## Questions this post answers

### What is OpenAI's Ultrafast mode for GPT-5.6 Sol and how much faster is it?

Ultrafast is a new OpenAI API tier that runs GPT-5.6 Sol up to 14 times faster than standard inference, reaching up to 750 output tokens per second. It uses Cerebras Wafer-Scale Engine chips, which keep 44 GB of SRAM on-chip to avoid the memory bandwidth bottlenecks typical GPU setups face. It is currently rolling out to a small group of customers as a limited preview.

_daily.dev tracks releases like this for teams evaluating inference speed as a factor in choosing AI providers._

### How does GPT-5.6 Sol on Ultrafast compare to Claude Fable 5 on Humanity's Last Exam?

GPT-5.6 Sol running on OpenAI's Ultrafast tier completed all 2,500 questions on Humanity's Last Exam in 11 hours and 11 minutes, compared to 78 hours and 27 minutes for Claude Fable 5 on the same benchmark. Cerebras also reports a 5.6x speedup on GDP-Val with no drop in output quality.

_engineers comparing model speed and cost tradeoffs can follow benchmark news like this on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 2 discussions and 107 comments across hackernews, x (as of 2026-08-19).

**TL;DR:** Discussion is mostly speculative, with people trying to reverse-engineer GPT-5.6 Sol's active parameter count from its speed rather than reacting to the Cerebras/Ultrafast announcement itself. Accelerating GPT-5.6 Sol Ultrafast with OpenAI: Developers are impressed by the raw speed and see real use cases (finance, live debugging, incident response), but many push back on the benchmark framing, question pricing/margins, and doubt claims of .

**Sentiment:** 38% positive · 42% mixed · 20% skeptical

**The case for**

- Speed is seen as genuinely valuable for latency-sensitive use cases like live analysis, trading, and incident debugging.
- Some argue faster iteration itself improves output quality by enabling more review/revision passes.
- The technical approach (wafer-scale on-chip SRAM avoiding batching) is seen as a clever, differentiated architecture bet.

**The pushback**

- Some see the speed as an inadvertent hint that the model has a much smaller active parameter count than expected, raising questions about its actual scale.
- Several question whether the HLE benchmark comparison is misleading since answering many independent questions is embarrassingly parallel, not a true measure of single-answer latency.
- Skepticism that 'no quality degradation' claims are trustworthy, citing past history of similar claims not holding up.
- Concerns about pricing/economics, expecting the service to be extremely expensive and only viable for well-funded customers.
- Some note that faster token generation doesn't fix other bottlenecks like test suites, typechecking, or grep, so real-world task speedup will be less than the raw token-rate gain.

**By community**

- hackernews (mixed): Excitement about the speed and architecture is tempered by pointed skepticism about the benchmark methodology, quality-parity claims, and likely high cost.
- x (mixed): Replies focus on guessing active parameter counts from throughput comparisons to other models rather than judging the Ultrafast tier itself.

**Hottest debate:** Whether the fast token throughput reveals that GPT-5.6 Sol has a surprisingly small active parameter count.

**Open questions**

- What is GPT-5.6 Sol's actual active parameter count?
- What will the actual pricing be for Ultrafast access?
- Why does a dense model like Gemini reportedly run fast without a routing mechanism?
- Is GPT-5.6 Sol truly quality-identical when run on Cerebras hardware, or is there hidden quantization/optimization tradeoff?
- Why can't Cerebras use its speed advantage for batching to serve more users instead of only lower latency?

**Highlights**

> @scaling01 wait is this not a huge sizedox for 5.6??
> — [januarycomputer on x · 8 points, 2 comments](https://x.com/januarycomputer/status/2089879418848366830)

> @januarycomputer @scaling01 let's see... GLM 4.7: 355B total params, 32B active: ~2,000 tps Kimi K2.7: 1T total, 32B active: ~2,000 TPS so active = main factor unfortunately, I think only the active param count can be inferred, not total 😔 emaybe active parameters < 100B for GPT-5.6 Sol?
> — [Algorithon on x · 4 points, 1 comments](https://x.com/Algorithon/status/2089886373419364517)

> @Algorithon @scaling01 naively this makes since since 70b models are about ~2x the 32b active crew. no idea why gemini dense is so fast though. maybe no routing?
> — [januarycomputer on x · 1 points](https://x.com/januarycomputer/status/2089887653621850283)

> Answering 2,500 independent questions is an embarrassingly parallel workload, all it needs is scale out.  It would be more meaningful to know how much time was required for a single complete answer to a difficult HLE question.
> — [zozbot234 on hackernews · 4 comments](https://news.ycombinator.com/item?id=49291102)

> Cerebras is a large plate sized chip. It has 50GB of SRAM, and few hundred K simple cores that can access that SRAM really fast. I don't know semiconductors well, but I understand that the same manufacturing technique that makes this huge chip possible, on the flip-side limits inter-chip communcation bandwidth. In cerebras, it is 150 GB/s (compared to nvlink's 2TB/s or groq's similar). One way large models are served on a bunch of cerebras chips is by essentially distributing layers' weights across chips. Few layers's weights per chip - as many as the KV cache + activations + weights will allow. You use pipelining to hide the latency of the inter-chip 150 GB/s link. On GPUs, you amortize the cost of loading weights from HBM to SRAM across multiple users - thereby making it cheaper _per_ user. But here, there is no such amortization. The weights are already there. It is the activations that stream through. You _could_ do batching/continuous batching, but that would just service more users at lower token/s without any amortization of fixed cost, due to fixed cost being non-existent.
> — [porridgeraisin on hackernews](https://news.ycombinator.com/item?id=49291477)

**Source threads**

- [hackernews](https://news.ycombinator.com/item?id=49289844) · 228 points · 100 comments
- [x](https://x.com/scaling01/status/2089875545488056322) · 0 points · 7 comments

## Similar posts on daily.dev

- [OpenAI unveils first model running on Cerebras silicon](https://daily.dev/posts/openai-unveils-first-model-running-on-cerebras-silicon-yyxkpjgbb) · The Register · 1 upvotes · 0 comments
- [OpenAI sidesteps Nvidia with unusually fast coding model on plate-sized chips](https://daily.dev/posts/openai-sidesteps-nvidia-with-unusually-fast-coding-model-on-plate-sized-chips-2dku3o06c) · Ars Technica · 2 upvotes · 0 comments
- [GPT-5.6 kernel of truth: Sol can cut its own costs, says OpenAI](https://daily.dev/posts/gpt-5-6-kernel-of-truth-sol-can-cut-its-own-costs-says-openai-q2dfgmfke) · The New Stack · 4 upvotes · 1 comments

---

Tags: [#ai](https://daily.dev/tags/ai), [#openai](https://daily.dev/tags/openai), [#ai-inference](https://daily.dev/tags/ai-inference)

[View this post on daily.dev](https://daily.dev/posts/openai-s-new-ultrafast-mode-runs-gpt-5-6-sol-14x-faster-using-cerebras-chips-kewuoa7pc)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"OpenAI's new Ultrafast mode runs GPT-5.6 Sol 14x faster using Cerebras chips","url":"https://daily.dev/posts/openai-s-new-ultrafast-mode-runs-gpt-5-6-sol-14x-faster-using-cerebras-chips-kewuoa7pc","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/openai-s-new-ultrafast-mode-runs-gpt-5-6-sol-14x-faster-using-cerebras-chips-kewuoa7pc"},"datePublished":"2026-08-13T20:07:38.773Z","dateModified":"2026-08-19T03:43:26.803Z","description":"OpenAI is previewing a new API tier called Ultrafast that runs GPT-5.6 Sol up to 14x faster than standard inference, reaching up to 750 output tokens per...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/fb36f5fda174e92fc9e8857e0947a2c3?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/fb36f5fda174e92fc9e8857e0947a2c3?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/openai-s-new-ultrafast-mode-runs-gpt-5-6-sol-14x-faster-using-cerebras-chips-kewuoa7pc","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,openai,ai-inference","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"OpenAI's new Ultrafast mode runs GPT-5.6 Sol 14x faster using Cerebras chips"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/openai-s-new-ultrafast-mode-runs-gpt-5-6-sol-14x-faster-using-cerebras-chips-kewuoa7pc#faq","mainEntity":[{"@type":"Question","name":"What is OpenAI's Ultrafast mode for GPT-5.6 Sol and how much faster is it?","acceptedAnswer":{"@type":"Answer","text":"Ultrafast is a new OpenAI API tier that runs GPT-5.6 Sol up to 14 times faster than standard inference, reaching up to 750 output tokens per second. It uses Cerebras Wafer-Scale Engine chips, which keep 44 GB of SRAM on-chip to avoid the memory bandwidth bottlenecks typical GPU setups face. It is currently rolling out to a small group of customers as a limited preview. daily.dev tracks releases like this for teams evaluating inference speed as a factor in choosing AI providers."}},{"@type":"Question","name":"How does GPT-5.6 Sol on Ultrafast compare to Claude Fable 5 on Humanity's Last Exam?","acceptedAnswer":{"@type":"Answer","text":"GPT-5.6 Sol running on OpenAI's Ultrafast tier completed all 2,500 questions on Humanity's Last Exam in 11 hours and 11 minutes, compared to 78 hours and 27 minutes for Claude Fable 5 on the same benchmark. Cerebras also reports a 5.6x speedup on GDP-Val with no drop in output quality. engineers comparing model speed and cost tradeoffs can follow benchmark news like this on daily.dev."}}]}
```

