<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/gpt-6-astra-benchmarks-deep-dive-this-is-not-a-good-coding-model-anymore---worse-than-fable--74nd09qsi" -->

---
title: GPT-6 Astra (Benchmarks Deep-dive): This is not a good...
description: A deep-dive analysis argues that OpenAI&#x27;s newly launched GPT6 Astra, despite marketing itself as the &#x27;world&#x27;s most intelligent model&#x27; ushering in the &#x27;AGI...
canonical: https://daily.dev/posts/gpt-6-astra-benchmarks-deep-dive-this-is-not-a-good-coding-model-anymore---worse-than-fable--74nd09qsi
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model anymore? - Worse than Fable? | daily.dev
og:description: A deep-dive analysis argues that OpenAI&#x27;s newly launched GPT6 Astra, despite marketing itself as the &#x27;world&#x27;s most intelligent model&#x27; ushering in the &#x27;AGI...
og:url: https://daily.dev/posts/gpt-6-astra-benchmarks-deep-dive-this-is-not-a-good-coding-model-anymore---worse-than-fable--74nd09qsi
og:image: https://api.daily.dev/og/posts/74nD09qsI.png
og:image:alt: GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model anymore? - Worse than Fable?
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model anymore? - Worse than Fable?

**[AICodeKing](https://daily.dev/sources/aicodeking)** · 14 min read · 0 upvotes · 0 comments

## Summary

A deep-dive analysis argues that OpenAI's newly launched GPT6 Astra, despite marketing itself as the 'world's most intelligent model' ushering in the 'AGI era,' is not a dominant coding upgrade. While Astra posts a striking 99.9% on ARC-AGI-3 (achieved using OpenAI's own provider adapter harness rather than the standard neutral one, where it scores 62.7%), its coding benchmarks (DeepSWE, Frontier Code, Artificial Analysis coding index) show it roughly tied with or behind Claude Fable 5.1 and only marginally ahead of GPT 5.6 Soul. Its broad intelligence index score of 61 ties Soul and trails Fable 5.1's 66, while pricing is 2.5x higher per token. Astra's real strengths appear to be computer use, terminal workflows, cybersecurity exploitation, and token-efficient long-horizon agentic tasks, while conventional coding, front-end polish, and price-to-performance lag behind competitors.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=qQzGm2-yVfM>

## Questions this post answers

### Is GPT6 Astra better than Claude Fable 5.1 at coding?

No, GPT6 Astra is roughly tied with or slightly behind Claude Fable 5.1 on most coding benchmarks. On the Artificial Analysis coding agent index, Fable 5.1 leads with 70 versus Astra's 67. On Frontier Code, Fable 5 scores 53.5 versus Astra's 53.3. Astra only clearly wins on Terminal Bench 4.0 (57.9% vs Fable's 55.8%) and uses about a third fewer tokens, making it more cost-efficient per task.

_Compare real coding benchmark deltas between frontier models before picking one for daily.dev readers building agent workflows._

### Why did GPT6 Astra score 99.9% on ARC-AGI-3 when it only scores 62.7% elsewhere?

The 99.9% score came from testing in OpenAI's own provider adapter harness, which preserves the model's private reasoning state between actions and uses OpenAI's compaction system, at a cost of about $18,800. In the standard provider-neutral harness, Astra scored only 62.7% at max reasoning, costing over $26,000. OpenAI's launch chart compared the 99.9% result against older models tested under different conditions, making it not a clean apples-to-apples comparison.

_Track how vendors report benchmark results differently when evaluating new coding models on daily.dev._

### How does GPT6 Astra's pricing compare to GPT 5.6 Soul?

GPT6 Astra costs $10 per million input tokens and $50 per million output tokens, compared to GPT 5.6 Soul's rate, making Astra roughly 2.5 times more expensive per token. Even though Astra uses about 10% fewer output tokens on the intelligence index, that efficiency gain does not offset the price hike, resulting in a total cost per task about 75% higher than Soul at max effort.

_Weigh cost-per-task tradeoffs across coding models by following pricing changes on daily.dev._

## Similar posts on daily.dev

- [GPT-6 Astra Benchmarks: 99.9% or 62.7%? The Full ARC-AGI-3 Story](https://daily.dev/posts/gpt-6-astra-benchmarks-99-9-or-62-7-the-full-arc-agi-3-story-1dabqghdx) · Medium · 1 upvotes · 0 comments

---

Tags: [#ai](https://daily.dev/tags/ai), [#ai-agents](https://daily.dev/tags/ai-agents), [#openai](https://daily.dev/tags/openai), [#claude](https://daily.dev/tags/claude)

[View this post on daily.dev](https://daily.dev/posts/gpt-6-astra-benchmarks-deep-dive-this-is-not-a-good-coding-model-anymore---worse-than-fable--74nd09qsi)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model anymore? - Worse than Fable?","url":"https://daily.dev/posts/gpt-6-astra-benchmarks-deep-dive-this-is-not-a-good-coding-model-anymore---worse-than-fable--74nd09qsi","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/gpt-6-astra-benchmarks-deep-dive-this-is-not-a-good-coding-model-anymore---worse-than-fable--74nd09qsi"},"datePublished":"2026-09-04T05:30:11.231Z","dateModified":"2026-09-04T22:00:50.455Z","description":"A deep-dive analysis argues that OpenAI's newly launched GPT6 Astra, despite marketing itself as the 'world's most intelligent model' ushering in the 'AGI...","image":"https://i.ytimg.com/vi/qQzGm2-yVfM/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/qQzGm2-yVfM/sddefault.jpg","isAccessibleForFree":true,"articleSection":"AICodeKing","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"AICodeKing","logo":"https://media.daily.dev/image/upload/s--x7nDUfWj--/f_auto,q_auto/v1768208312/logos/aicodeking","url":"https://daily.dev/sources/aicodeking"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/gpt-6-astra-benchmarks-deep-dive-this-is-not-a-good-coding-model-anymore---worse-than-fable--74nd09qsi","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,ai-agents,openai,claude","timeRequired":"PT14M","video":{"@type":"VideoObject","name":"GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model anymore? - Worse than Fable?","description":"A deep-dive analysis argues that OpenAI's newly launched GPT6 Astra, despite marketing itself as the 'world's most intelligent model' ushering in the 'AGI...","thumbnailUrl":"https://i.ytimg.com/vi/qQzGm2-yVfM/sddefault.jpg","uploadDate":"2026-09-04T05:30:11.231Z","duration":"PT14M","url":"https://api.daily.dev/r/74nD09qsI","embedUrl":"https://www.youtube.com/embed/qQzGm2-yVfM"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"AICodeKing","item":"https://daily.dev/sources/aicodeking"},{"@type":"ListItem","position":3,"name":"GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model anymore? - Worse than Fable?"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/gpt-6-astra-benchmarks-deep-dive-this-is-not-a-good-coding-model-anymore---worse-than-fable--74nd09qsi#faq","mainEntity":[{"@type":"Question","name":"Is GPT6 Astra better than Claude Fable 5.1 at coding?","acceptedAnswer":{"@type":"Answer","text":"No, GPT6 Astra is roughly tied with or slightly behind Claude Fable 5.1 on most coding benchmarks. On the Artificial Analysis coding agent index, Fable 5.1 leads with 70 versus Astra's 67. On Frontier Code, Fable 5 scores 53.5 versus Astra's 53.3. Astra only clearly wins on Terminal Bench 4.0 (57.9% vs Fable's 55.8%) and uses about a third fewer tokens, making it more cost-efficient per task. Compare real coding benchmark deltas between frontier models before picking one for daily.dev readers building agent workflows."}},{"@type":"Question","name":"Why did GPT6 Astra score 99.9% on ARC-AGI-3 when it only scores 62.7% elsewhere?","acceptedAnswer":{"@type":"Answer","text":"The 99.9% score came from testing in OpenAI's own provider adapter harness, which preserves the model's private reasoning state between actions and uses OpenAI's compaction system, at a cost of about $18,800. In the standard provider-neutral harness, Astra scored only 62.7% at max reasoning, costing over $26,000. OpenAI's launch chart compared the 99.9% result against older models tested under different conditions, making it not a clean apples-to-apples comparison. Track how vendors report benchmark results differently when evaluating new coding models on daily.dev."}},{"@type":"Question","name":"How does GPT6 Astra's pricing compare to GPT 5.6 Soul?","acceptedAnswer":{"@type":"Answer","text":"GPT6 Astra costs $10 per million input tokens and $50 per million output tokens, compared to GPT 5.6 Soul's rate, making Astra roughly 2.5 times more expensive per token. Even though Astra uses about 10% fewer output tokens on the intelligence index, that efficiency gain does not offset the price hike, resulting in a total cost per task about 75% higher than Soul at max effort. Weigh cost-per-task tradeoffs across coding models by following pricing changes on daily.dev."}}]}
```

