<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/meta-releases-muse-spark-1-3-claims-top-scores-on-deepswe-benchmark-sitagwxqs" -->

---
title: Meta releases Muse Spark 1.3, claims top scores on...
description: Meta has begun rolling out Muse Spark 1.3, available through Muse Code and the Meta model API. Meta calls it the company&#x27;s biggest improvement yet on coding...
canonical: https://daily.dev/posts/meta-releases-muse-spark-1-3-claims-top-scores-on-deepswe-benchmark-sitagwxqs
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Meta releases Muse Spark 1.3, claims top scores on DeepSWE benchmark | daily.dev
og:description: Meta has begun rolling out Muse Spark 1.3, available through Muse Code and the Meta model API. Meta calls it the company&#x27;s biggest improvement yet on coding...
og:url: https://daily.dev/posts/meta-releases-muse-spark-1-3-claims-top-scores-on-deepswe-benchmark-sitagwxqs
og:image: https://api.daily.dev/og/posts/sitAgWxqS.png
og:image:alt: Meta releases Muse Spark 1.3, claims top scores on DeepSWE benchmark
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Meta releases Muse Spark 1.3, claims top scores on DeepSWE benchmark

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

Meta has begun rolling out Muse Spark 1.3, available through Muse Code and the Meta model API. Meta calls it the company's biggest improvement yet on coding and agentic tasks, claiming it outscores GPT-5.6 and Opus 5 on the DeepSWE 1.1 benchmark, though independent verification is pending. Pricing is said to be significantly cheaper than comparable frontier models, but no numbers have been published. Meta also teased two upcoming releases: an open-weights version of Muse Spark and a larger model codenamed Watermelon.

## Content

Meta has released Muse Spark 1.3, which the company's Chief AI Officer Alexandr Wang called its biggest performance jump yet. The model is rolling out now through Meta's API and in Muse Code.

## Benchmark performance

On the Artificial Analysis Intelligence Index, Muse Spark 1.3 now ranks as the top non-Anthropic model, matching Claude Fable 5 and sitting just one point behind Opus 5 and four points behind Fable 5.1. On DeepSWE 1.1, it scores above both GPT-5.6 and Opus 5.

Long-context performance is a particular strength: the model scores 98.5 on MRCR 256K–512K, beating GPT-5.6 Sol. On agentic benchmarks it gets close to Opus 5 — JobBench (64.9 vs 65.7), OSWorld (66.9 vs 68.3), and AutomationBench (49.4 vs 50.3).

Compared to Muse Spark 1.2, the new version uses 20% fewer tool calls and 25% fewer tokens to complete the same tasks.

## The cost gap is hard to ignore

Output tokens cost $4.25 per million, versus $50 per million for Fable 5. On a per-task basis, the extended ("xhigh") configuration runs $0.55 compared to $3.14 for Fable — about 5.7x cheaper. Whether a few benchmark points justify that price difference is a reasonable question, and one that gets harder to answer as models like Muse Spark close the gap.

## Safety changes

Meta made several agentic safety improvements in this release: the model is better at recognizing its own limits, asks for confirmation before taking irreversible actions, and generates fewer tokens per task. These changes follow an earlier incident where Muse Spark 1.1 accessed an external service during testing — later traced to a vendor that left evaluation environments exposed.

## Open weights: still undecided for 1.3

Meta hasn't decided whether to release the weights for 1.3. The company still plans to release 1.2's weights, and the first Muse Spark launched closed source. The decision has regulatory implications: under Article 53 of the EU AI Act, genuinely open-source general-purpose models are exempt from certain technical documentation requirements — but that exemption disappears entirely for models classified as carrying systemic risk, regardless of license.

Meta also teased two upcoming releases: "Watermelon," described as the next major model upgrade, and an open-weight version of Muse Spark.

## Questions this post answers

### What is Meta Muse Spark 1.3 and how does it compare to GPT-5.6 and Opus 5?

Muse Spark 1.3 is Meta's latest coding and agentic model, rolled out through Muse Code and the Meta model API. Meta claims it scores above GPT-5.6 and Opus 5 on the DeepSWE 1.1 benchmark, calling it the company's biggest improvement yet on coding and agentic tasks, though independent verification of the benchmark result is still pending.

_Developers picking a coding model can track how Muse Spark's benchmark claims hold up on daily.dev._

### Is Meta releasing an open-weights version of Muse Spark?

Yes, Meta has confirmed an open-weights version of Muse Spark is in its release pipeline, alongside a separate larger model internally codenamed 'Watermelon.' Neither has a confirmed release date or specifications yet, and pricing details for Muse Spark 1.3 itself also have not been published.

_Anyone weighing open versus closed coding models can follow Muse Spark's rollout on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 3 discussions and 32 comments across x (as of 2026-09-03).

**TL;DR:** Reaction centers on the aggressive price-to-performance ratio versus Anthropic and OpenAI, with genuine enthusiasm about efficiency gains tempered by skepticism that benchmark scores translate to real-world code quality.

**Sentiment:** 45% positive · 40% mixed · 15% skeptical

**The case for**

- The output token price ($4.25/M vs $50/M) for near-parity scores is seen as a major disruption to incumbent pricing.
- Some hands-on users report strong, fast performance after real usage.
- Efficiency gains (fewer tool calls/tokens) are seen as the more meaningful story for long-horizon agentic tasks than raw benchmark score.

**The pushback**

- Doubts that benchmark scores reflect real frontend polish or mergeable, maintainable code quality.
- Concern that per-task cost advantage still needs more data on higher configurations to confirm it holds.
- Some argue OpenAI and Anthropic models remain ahead on other benchmarks like cyber, AI research, and ARC-AGI despite this index.
- One user found the model's output quality subjectively 'off' after trying it.

**By community**

- x (mixed): Excitement about the price disruption and efficiency gains is countered by skepticism over whether benchmark wins reflect real coding quality.

**Hottest debate:** Whether the headline benchmark and price advantages are decision-relevant or just a superficial 'benchmark contest' that ignores real-world code quality.

**Open questions**

- Does the per-task cost advantage hold up at higher (max) configurations with more data?
- How does Muse 1.3 actually perform on messy real-world codebases versus benchmark tasks?
- How does it compare against other open models like GLM or Kimi?

**Highlights**

> @Hesamation Same score as Fable 5 but $4.25/M vs $50/M means Opus is the overpriced bottle at a beer festival. buy speed every time
> — [TheAIShrink on x](https://x.com/TheAIShrink/status/2095281834942714132)

> @tobi The more I see it, the more I stop believing these numbers. Maybe it's because I also check front-end designs and mergeable code quality. I feel like these smaller agents would execute good stuff big agents plan. I think these tests don't really check that kind of things.
> — [M\_I\_H\_111 on x · 1 points, 2 comments](https://x.com/M_I_H_111/status/2095284284454211741)

> @Hesamation the xhigh result is more decision-relevant than the token-price headline max still needs per-task data to confirm that advantage holds
> — [EgorJioo on x](https://x.com/EgorJioo/status/2095279132632379873)

> @Hesamation OpenAi & Anthropic models are miles ahead in Cyber and Ai Research (as well as ArcAGI, METR time horizons, etc.) even if AAII says ow
> — [Hong60282445 on x](https://x.com/Hong60282445/status/2095301785036468413)

> @tobi That Muse 1.3 benchmark run is genuinely impressive. Still waiting to see how it handles messy real-world code though.
> — [websterweby on x · 1 points](https://x.com/websterweby/status/2095286733512528371)

**Source threads**

- [x](https://x.com/Hesamation/status/2095276123479318914) · 0 points · 9 comments
- [x](https://x.com/tobi/status/2095279970058731549) · 0 points · 23 comments
- [x](https://x.com/yacineMTB/status/2095315317539078355) · 0 points · 0 comments

## Similar posts on daily.dev

- [Meta enters the crowded AI coding battle with Muse Spark 1.1](https://daily.dev/posts/meta-enters-the-crowded-ai-coding-battle-with-muse-spark-1-1-wtfjrleyt) · TechCrunch · 1 upvotes · 0 comments

---

Tags: [#ai](https://daily.dev/tags/ai), [#ai-coding](https://daily.dev/tags/ai-coding)

[View this post on daily.dev](https://daily.dev/posts/meta-releases-muse-spark-1-3-claims-top-scores-on-deepswe-benchmark-sitagwxqs)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Meta releases Muse Spark 1.3, claims top scores on DeepSWE benchmark","url":"https://daily.dev/posts/meta-releases-muse-spark-1-3-claims-top-scores-on-deepswe-benchmark-sitagwxqs","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/meta-releases-muse-spark-1-3-claims-top-scores-on-deepswe-benchmark-sitagwxqs"},"datePublished":"2026-09-02T19:34:12.630Z","dateModified":"2026-09-03T01:36:05.445Z","description":"Meta has begun rolling out Muse Spark 1.3, available through Muse Code and the Meta model API. Meta calls it the company's biggest improvement yet on coding...","image":"https://pbs.twimg.com/media/HRPD4HNacAA-i4R.png","thumbnailUrl":"https://pbs.twimg.com/media/HRPD4HNacAA-i4R.png","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/meta-releases-muse-spark-1-3-claims-top-scores-on-deepswe-benchmark-sitagwxqs","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,ai-coding","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Meta releases Muse Spark 1.3, claims top scores on DeepSWE benchmark"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/meta-releases-muse-spark-1-3-claims-top-scores-on-deepswe-benchmark-sitagwxqs#faq","mainEntity":[{"@type":"Question","name":"What is Meta Muse Spark 1.3 and how does it compare to GPT-5.6 and Opus 5?","acceptedAnswer":{"@type":"Answer","text":"Muse Spark 1.3 is Meta's latest coding and agentic model, rolled out through Muse Code and the Meta model API. Meta claims it scores above GPT-5.6 and Opus 5 on the DeepSWE 1.1 benchmark, calling it the company's biggest improvement yet on coding and agentic tasks, though independent verification of the benchmark result is still pending. Developers picking a coding model can track how Muse Spark's benchmark claims hold up on daily.dev."}},{"@type":"Question","name":"Is Meta releasing an open-weights version of Muse Spark?","acceptedAnswer":{"@type":"Answer","text":"Yes, Meta has confirmed an open-weights version of Muse Spark is in its release pipeline, alongside a separate larger model internally codenamed 'Watermelon.' Neither has a confirmed release date or specifications yet, and pricing details for Muse Spark 1.3 itself also have not been published. Anyone weighing open versus closed coding models can follow Muse Spark's rollout on daily.dev."}}]}
```

