Collection

OpenAI's new Ultrafast mode runs GPT-5.6 Sol 14x faster using Cerebras chips

6 sources
Post cover image

Questions this post answers

What is OpenAI's Ultrafast mode for GPT-5.6 Sol and how much faster is it?

Ultrafast is a new OpenAI API tier that runs GPT-5.6 Sol up to 14 times faster than standard inference, reaching up to 750 output tokens per second. It uses Cerebras Wafer-Scale Engine chips, which keep 44 GB of SRAM on-chip to avoid the memory bandwidth bottlenecks typical GPU setups face. It is currently rolling out to a small group of customers as a limited preview. daily.dev tracks releases like this for teams evaluating inference speed as a factor in choosing AI providers.

How does GPT-5.6 Sol on Ultrafast compare to Claude Fable 5 on Humanity's Last Exam?

GPT-5.6 Sol running on OpenAI's Ultrafast tier completed all 2,500 questions on Humanity's Last Exam in 11 hours and 11 minutes, compared to 78 hours and 27 minutes for Claude Fable 5 on the same benchmark. Cerebras also reports a 5.6x speedup on GDP-Val with no drop in output quality. engineers comparing model speed and cost tradeoffs can follow benchmark news like this on daily.dev.

80 Impressions