<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/i-love-ultrafast-it-s-unusable--c8foalwuy" -->

---
title: I love Ultrafast (it's unusable) | daily.dev
description: A creator reviews OpenAI's new 'Ultrafast' inference mode for the Astra model inside Codex, which generates code at extremely high token-per-second speeds...
canonical: https://daily.dev/posts/i-love-ultrafast-it-s-unusable--c8foalwuy
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: I love Ultrafast (it's unusable) | daily.dev
og:description: A creator reviews OpenAI's new 'Ultrafast' inference mode for the Astra model inside Codex, which generates code at extremely high token-per-second speeds...
og:url: https://daily.dev/posts/i-love-ultrafast-it-s-unusable--c8foalwuy
og:image: https://api.daily.dev/og/posts/C8foAlWuY.png
og:image:alt: I love Ultrafast (it's unusable)
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# I love Ultrafast (it's unusable)

**[Theo - t3․gg](https://daily.dev/sources/t3dotgg)** · 28 min read · 0 upvotes · 0 comments

## Summary

A creator reviews OpenAI's new 'Ultrafast' inference mode for the Astra model inside Codex, which generates code at extremely high token-per-second speeds (over 300 TPS) but at drastically inflated prices: $300-$450 per million output tokens versus $50 normally, with cache writes at $75/mill. Testing it on small pull request reviews cost $600, and building a side project (Slopolytics) cost $286 versus roughly $12 using Sonnet. Despite the brutal economics, the author argues the real-time, in-the-loop prompting experience changes how you work and iterate on UI, though he concludes almost nobody should pay for it at current prices, especially the $500/month plan. He notes OpenAI employees reportedly now need approval to use it due to compute costs, and that a future 'Sonnet Ultrafast' mode at similar multipliers could make the speed tier actually worth it.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=pJljViiUEPw>

## Questions this post answers

### How much more expensive is OpenAI's Ultrafast mode compared to standard Astra pricing?

Ultrafast inference costs roughly 6 times more than standard Astra pricing, rising from $50 to $300 per million output tokens, and up to $450 per million tokens for long-context projects. Cache write costs jump to $75 per million tokens and cache reads to $6 per million, both far above normal rates. A side project that cost $12 with Sonnet cost $286 with Astra on Ultrafast.

_Developers weighing speed against budget on AI coding tools can track real pricing shifts like this through daily.dev._

### Is it worth paying $500 a month for OpenAI's Ultrafast Codex plan?

The $500 per month Ultrafast plan is generally not worth it, since even with roughly $1,000 of subsidized usage per week, that allowance can be burned through in about 2.1 hours of active use. OpenAI employees reportedly now need approval per task to use it due to its cost, and the recommendation is to avoid the tier until a cheaper high-speed option for other models arrives.

_Anyone budgeting for premium AI coding subscriptions can follow real cost breakdowns like this via daily.dev before upgrading._

### What is the token throughput difference between regular, fast, and ultra fast inference modes in Codex?

Regular mode runs around 30 tokens per second, fast mode around 60 tokens per second, and ultra fast mode exceeds 300 tokens per second, typically between 320 and 340 TPS. This roughly tenfold speed increase over regular mode enables real-time, live-updating code generation but comes with a proportionally extreme price increase.

_Developers comparing inference speed tiers across coding assistants can keep tabs on throughput benchmarks through daily.dev._

## Similar posts on daily.dev

- [Google's Gemini 3.5 Flash costs 3x the model it replaced, and the era of cheap AI is ending](https://daily.dev/posts/google-s-gemini-3-5-flash-costs-3x-the-model-it-replaced-and-the-era-of-cheap-ai-is-ending-ochkeykzb) · XDA Developers · 2 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-coding](https://daily.dev/tags/ai-coding), [#vibe-coding](https://daily.dev/tags/vibe-coding), [#openai-codex](https://daily.dev/tags/openai-codex)

[View this post on daily.dev](https://daily.dev/posts/i-love-ultrafast-it-s-unusable--c8foalwuy)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"I love Ultrafast (it's unusable)","url":"https://daily.dev/posts/i-love-ultrafast-it-s-unusable--c8foalwuy","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/i-love-ultrafast-it-s-unusable--c8foalwuy"},"datePublished":"2026-10-06T13:08:09.012Z","dateModified":"2026-10-06T13:19:21.750Z","description":"A creator reviews OpenAI's new 'Ultrafast' inference mode for the Astra model inside Codex, which generates code at extremely high token-per-second speeds...","image":"https://i.ytimg.com/vi/pJljViiUEPw/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/pJljViiUEPw/sddefault.jpg","isAccessibleForFree":true,"articleSection":"Theo - t3․gg","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Theo - t3․gg","logo":"https://media.daily.dev/image/upload/s--UmX7IyU3--/f_auto/v1704628081/logos/t3dotgg.jpg","url":"https://daily.dev/sources/t3dotgg"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/i-love-ultrafast-it-s-unusable--c8foalwuy","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-coding,vibe-coding,openai-codex","timeRequired":"PT28M","video":{"@type":"VideoObject","name":"I love Ultrafast (it's unusable)","description":"A creator reviews OpenAI's new 'Ultrafast' inference mode for the Astra model inside Codex, which generates code at extremely high token-per-second speeds...","thumbnailUrl":"https://i.ytimg.com/vi/pJljViiUEPw/sddefault.jpg","uploadDate":"2026-10-06T13:08:09.012Z","duration":"PT28M","url":"https://api.daily.dev/r/C8foAlWuY","embedUrl":"https://www.youtube.com/embed/pJljViiUEPw"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Theo - t3․gg","item":"https://daily.dev/sources/t3dotgg"},{"@type":"ListItem","position":3,"name":"I love Ultrafast (it's unusable)"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/i-love-ultrafast-it-s-unusable--c8foalwuy#faq","mainEntity":[{"@type":"Question","name":"How much more expensive is OpenAI's Ultrafast mode compared to standard Astra pricing?","acceptedAnswer":{"@type":"Answer","text":"Ultrafast inference costs roughly 6 times more than standard Astra pricing, rising from $50 to $300 per million output tokens, and up to $450 per million tokens for long-context projects. Cache write costs jump to $75 per million tokens and cache reads to $6 per million, both far above normal rates. A side project that cost $12 with Sonnet cost $286 with Astra on Ultrafast. Developers weighing speed against budget on AI coding tools can track real pricing shifts like this through daily.dev."}},{"@type":"Question","name":"Is it worth paying $500 a month for OpenAI's Ultrafast Codex plan?","acceptedAnswer":{"@type":"Answer","text":"The $500 per month Ultrafast plan is generally not worth it, since even with roughly $1,000 of subsidized usage per week, that allowance can be burned through in about 2.1 hours of active use. OpenAI employees reportedly now need approval per task to use it due to its cost, and the recommendation is to avoid the tier until a cheaper high-speed option for other models arrives. Anyone budgeting for premium AI coding subscriptions can follow real cost breakdowns like this via daily.dev before upgrading."}},{"@type":"Question","name":"What is the token throughput difference between regular, fast, and ultra fast inference modes in Codex?","acceptedAnswer":{"@type":"Answer","text":"Regular mode runs around 30 tokens per second, fast mode around 60 tokens per second, and ultra fast mode exceeds 300 tokens per second, typically between 320 and 340 TPS. This roughly tenfold speed increase over regular mode enables real-time, live-updating code generation but comes with a proportionally extreme price increase. Developers comparing inference speed tiers across coding assistants can keep tabs on throughput benchmarks through daily.dev."}}]}
```

