<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/cognition-s-swe-2-matches-frontier-coding-benchmarks-at-64-lower-cost-agchp1hyh" -->

---
title: Cognition&#x27;s SWE-2 matches frontier coding benchmarks at...
description: Cognition released SWE-2, a specialized coding model scoring 50.0% on FrontierCode 1.1, within one point of a Fable 5.1 model, while costing 64% less. It&#x27;s now...
canonical: https://daily.dev/posts/cognition-s-swe-2-matches-frontier-coding-benchmarks-at-64-lower-cost-agchp1hyh
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Cognition&#x27;s SWE-2 matches frontier coding benchmarks at 64% lower cost | daily.dev
og:description: Cognition released SWE-2, a specialized coding model scoring 50.0% on FrontierCode 1.1, within one point of a Fable 5.1 model, while costing 64% less. It&#x27;s now...
og:url: https://daily.dev/posts/cognition-s-swe-2-matches-frontier-coding-benchmarks-at-64-lower-cost-agchp1hyh
og:image: https://api.daily.dev/og/posts/aGChp1HYh.png
og:image:alt: Cognition&#x27;s SWE-2 matches frontier coding benchmarks at 64% lower cost
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Cognition's SWE-2 matches frontier coding benchmarks at 64% lower cost

**[Trends](https://daily.dev/sources/trends)** · 2 min read · 0 upvotes · 2 comments

## Summary

Cognition released SWE-2, a specialized coding model scoring 50.0% on FrontierCode 1.1, within one point of a Fable 5.1 model, while costing 64% less. It's now live in Devin. Cognition claims RL scaling to multiple trillions of parameters with a refined training recipe. Reaction from ML commentators has been positive, framing it as specialized models pushing the Pareto frontier of cost versus capability. The piece also notes broader industry framing that frontier AI development is consolidating around two labs (implied Google and Anthropic), while specialized models like SWE-2 challenge the assumption that developers need the most expensive general-purpose model for coding tasks.

## Content

Cognition just released SWE-2, a coding model that scores 50.0% on FrontierCode 1.1 — within one point of Fable 5.1 — while costing 64% less. That's the whole pitch, and it's a pretty good one.

The technical claim: they scaled reinforcement learning to multiple trillions of parameters with a "refined recipe" that moves the Pareto curve on both capability and cost simultaneously. Scott Wu (Cognition) called it their first model with "performance comparable to the frontier" and is making it free for existing Devin plan users for the next month.

The reception is warm but measured. Kent C. Dodds and Omar Sarsur are enthusiastic about the cost-efficiency angle — the "specialized models pushing the Pareto frontier" framing is doing a lot of work in the replies. Nobody's calling it a GPT-killer. The honest read is that SWE-2 is competitive with top models on coding benchmarks, not ahead of them, and the real story is the price gap.

The broader context worth noting: Ethan Mollick observed this week that frontier AI is increasingly a two-company race, with Google and one other pulling ahead while everyone else clusters behind. SWE-2 lands in that second cluster — impressive, genuinely useful, but not a disruption to the top of the leaderboard. Cognition seems fine with that framing. "Closest to the frontier" is honest positioning.

The giveaway campaign (50 free $200 Max plans, influencer codes, reply-to-win mechanics) is running alongside the launch, which is either smart distribution or a sign they need to juice adoption. Probably both. Either way, if you're on a Devin plan, the free month is worth poking at.

## Questions this post answers

### What is Cognition's SWE-2 model and how does it compare to frontier models on coding benchmarks?

SWE-2 is a specialized coding model from Cognition that scores 50.0% on FrontierCode 1.1, within one point of Fable 5.1, while costing 64% less to run. It is now live in Devin. Cognition says they scaled reinforcement learning to multiple trillions of parameters with a refined training recipe to achieve this result.

_Developers weighing model cost against coding capability can track releases like this one on daily.dev._

### Is SWE-2 available for use yet?

Yes, SWE-2 is already live in Devin, Cognition's AI coding agent product, following its release with benchmark results showing near-frontier coding performance at a fraction of the cost of general frontier models.

_Anyone tracking which coding agents are shipping new models can follow updates like this on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 3 discussions and 129 comments across x (as of 2026-09-11).

**TL;DR:** Reaction is largely enthusiastic hype and giveaway participation, with a few users reporting confusing rate-limit issues that got resolved, and a rare skeptical voice framing the low-cost angle as the real story worth testing on longer tasks.

**Sentiment:** 55% positive · 35% mixed · 10% skeptical

**The case for**

- Some praise the model as highly capable and fast, including for zero-shot frontend work.
- One commenter frames the low per-rollout cost as the actual noteworthy result, more than the raw benchmark score.

**The pushback**

- Multiple users hit confusing usage limits shortly after starting, with unclear messaging about what counted toward quota.
- One person questioned whether the lower cost holds up on longer-running tasks rather than short benchmarks.

**By community**

- x (positive): Mostly congratulatory replies and giveaway entries, with a couple of resolved complaints about unclear usage limits and one comment probing whether the cost advantage holds on longer tasks.

**Open questions**

- Whether the touted cost/performance gap persists on longer-running agentic tasks rather than short benchmark rollouts.

**Highlights**

> @ScottWu46 the x-axis is the real story. near-fable scores at about a third of the cost per rollout. once agents are paying for their own tools, that's the number they'll shop on. does the gap hold on longer tasks?
> — [softaxiom\_ on x · 3 points](https://x.com/softaxiom_/status/2098235315722826162)

> @ScottWu46 It says my limits were exceeded and stopped working... Very misleading and dislike the ambiguity. I'm using Swe 2 and basically stopped working after 30 minutes.
> — [gdgjv18270 on x · 1 points, 1 comments](https://x.com/gdgjv18270/status/2098214608419033270)

> @ScottWu46 resolved! @dabit3  checked it out and I was using Cloud threads which don’t fall under the unlimited usage for SWE 2. Only applicable to local CLI/Desktop threads which is still super generous :)
> — [gdgjv18270 on x · 2 points, 1 comments](https://x.com/gdgjv18270/status/2098269960657305732)

> @ScottWu46 It honestly feels like Fable did when it first dropped. Shots fired...
> — [TannerPowell on x](https://x.com/TannerPowell/status/2098261961888985271)

**Source threads**

- [x](https://x.com/kloss_xyz/status/2098302482506273261) · 0 points · 0 comments
- [x](https://x.com/ScottWu46/status/2098147108771762641) · 0 points · 29 comments
- [x](https://x.com/ryancarson/status/2098147499542507662) · 0 points · 100 comments

## Community discussion

Top comments from developers on daily.dev.

**@davidbandel** · 1 upvotes

> swe-2 is heavily benchmaxed and possibly distilled. the instant it encountered a bench outside its distribution (terminal bench 4) it, along with the other benchmaxed, heavily distilled, models grok 4.6 & k3, failed spectacularly. meanwhile 5.6-Sol, Astra, and Fable cooked hard at it (at least compared to the rest of the pack)
>
> therefore swe-2 is either quantized, has a small parameter count (explaining its purported low cast/high inference speed), or is heavily distilled off of some other frontier model.
>
> i hate benchmaxed models. literally any benchmaxed model is inevitably trash at real...

**@geniusatwork** · 0 upvotes

> B
>
> ![image.png](https://media.daily.dev/image/upload/s--sghhbf3R--/f_auto/v1789106438/ugc/content_0a29031d-4662-4f40-8c3d-1d9407f246d6?_a=BAMAMicg0)
>
> Interesting. I'm on team skeptical.

---

Tags: [#ai](https://daily.dev/tags/ai), [#ai-coding](https://daily.dev/tags/ai-coding), [#finops](https://daily.dev/tags/finops), [#devin](https://daily.dev/tags/devin)

[View this post on daily.dev](https://daily.dev/posts/cognition-s-swe-2-matches-frontier-coding-benchmarks-at-64-lower-cost-agchp1hyh)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Cognition's SWE-2 matches frontier coding benchmarks at 64% lower cost","url":"https://daily.dev/posts/cognition-s-swe-2-matches-frontier-coding-benchmarks-at-64-lower-cost-agchp1hyh","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/cognition-s-swe-2-matches-frontier-coding-benchmarks-at-64-lower-cost-agchp1hyh"},"datePublished":"2026-09-10T16:00:58.893Z","dateModified":"2026-09-11T06:48:52.868Z","description":"Cognition released SWE-2, a specialized coding model scoring 50.0% on FrontierCode 1.1, within one point of a Fable 5.1 model, while costing 64% less. It's now...","image":"https://pbs.twimg.com/media/HR3fNU1bkAAchST.png","thumbnailUrl":"https://pbs.twimg.com/media/HR3fNU1bkAAchST.png","isAccessibleForFree":true,"articleSection":"Trends","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Trends","logo":"https://media.daily.dev/image/upload/s--ZfSp3asX--/f_auto,q_auto/v1780996004/logos/trends?_a=BAMAMiWQ0","url":"https://daily.dev/sources/trends"},"commentCount":2,"discussionUrl":"https://daily.dev/posts/cognition-s-swe-2-matches-frontier-coding-benchmarks-at-64-lower-cost-agchp1hyh","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":2}],"keywords":"ai,ai-coding,finops,devin","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Trends","item":"https://daily.dev/sources/trends"},{"@type":"ListItem","position":3,"name":"Cognition's SWE-2 matches frontier coding benchmarks at 64% lower cost"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/cognition-s-swe-2-matches-frontier-coding-benchmarks-at-64-lower-cost-agchp1hyh","comment":[{"@type":"Comment","text":"swe-2 is heavily benchmaxed and possibly distilled. the instant it encountered a bench outside its distribution (terminal bench 4) it, along with the other benchmaxed, heavily distilled, models grok 4.6 &amp; k3, failed spectacularly. meanwhile 5.6-Sol, Astra, and Fable cooked hard at it (at least compared to the rest of the pack)\ntherefore swe-2 is either quantized, has a small parameter count (explaining its purported low cast/high inference speed), or is heavily distilled off of some other frontier model.\ni hate benchmaxed models. literally any benchmaxed model is inevitably trash at real world work, long horizon tasks, and IF.\nin short, not impressed. just another slop model.","datePublished":"2026-09-10T19:03:13.393Z","url":"https://daily.dev/posts/aGChp1HYh#c-EcUlwAh8m","author":{"@type":"Person","name":"David Bandel","url":"https://daily.dev/davidbandel","image":"https://lh3.googleusercontent.com/a/ACg8ocL-0RYRxKo30FsvCz2FeR_E2FMAfOQ5C7SiJ2TTJuTt_qZnLA=s96-c"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"B\n\nInteresting. I’m on team skeptical.","datePublished":"2026-09-11T06:01:21.051Z","url":"https://daily.dev/posts/aGChp1HYh#c-yFc6aWR7g","author":{"@type":"Person","name":"Sam Dolin","url":"https://daily.dev/geniusatwork","image":"https://media.daily.dev/image/upload/s--nUhoZxKm--/f_auto/v1787434895/avatars/avatar_75QyPrL3BZgArSo6aFWuZ?_a=BAMAMicg0"}}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/cognition-s-swe-2-matches-frontier-coding-benchmarks-at-64-lower-cost-agchp1hyh#faq","mainEntity":[{"@type":"Question","name":"What is Cognition's SWE-2 model and how does it compare to frontier models on coding benchmarks?","acceptedAnswer":{"@type":"Answer","text":"SWE-2 is a specialized coding model from Cognition that scores 50.0% on FrontierCode 1.1, within one point of Fable 5.1, while costing 64% less to run. It is now live in Devin. Cognition says they scaled reinforcement learning to multiple trillions of parameters with a refined training recipe to achieve this result. Developers weighing model cost against coding capability can track releases like this one on daily.dev."}},{"@type":"Question","name":"Is SWE-2 available for use yet?","acceptedAnswer":{"@type":"Answer","text":"Yes, SWE-2 is already live in Devin, Cognition's AI coding agent product, following its release with benchmark results showing near-frontier coding performance at a fraction of the cost of general frontier models. Anyone tracking which coding agents are shipping new models can follow updates like this on daily.dev."}}]}
```

