<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/meta-muse-spark-1-2-benchmarks-cloudflare-rewrites-the-agentic-web-stack-gabz2pamn" -->

---
title: Meta Muse Spark 1.2 benchmarks, Cloudflare rewrites the...
description: Meta released Muse Spark 1.2 and Muse Code beta today, with benchmark scores that put it near GPT-5.5 territory at a fraction of the cost — though the...
canonical: https://daily.dev/posts/meta-muse-spark-1-2-benchmarks-cloudflare-rewrites-the-agentic-web-stack-gabz2pamn
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Meta Muse Spark 1.2 benchmarks, Cloudflare rewrites the agentic web stack | daily.dev
og:description: Meta released Muse Spark 1.2 and Muse Code beta today, with benchmark scores that put it near GPT-5.5 territory at a fraction of the cost — though the...
og:url: https://daily.dev/posts/meta-muse-spark-1-2-benchmarks-cloudflare-rewrites-the-agentic-web-stack-gabz2pamn
og:image: https://api.daily.dev/og/posts/gabZ2pamN.png
og:image:alt: Meta Muse Spark 1.2 benchmarks, Cloudflare rewrites the agentic web stack
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Meta Muse Spark 1.2 benchmarks, Cloudflare rewrites the agentic web stack

**[Agentic Digest](https://daily.dev/sources/agents_digest)** · 6 min read · 3 upvotes · 0 comments

## Summary

Meta released Muse Spark 1.2 and Muse Code beta today, with benchmark scores that put it near GPT-5.5 territory at a fraction of the cost — though the Contributor pricing tier's data-use terms deserve a close read before teams adopt it. Cloudflare dropped a coordinated set of releases rewriting how agents interact with the web: a stateless MCP spec, a V8-native headless browser, WebMCP for any hosted site, and an agent readiness scoring tool. Qwen3.8-Max open weights land next week, and a ChatGPT sandbox attack chain presented at Black Hat is worth understanding even though OpenAI patched the specific vector. DeepSeek signaled a significant price hike with no timeline, which matters given how many teams built cost assumptions around its rates.

## Content

**TLDR:** Meta released Muse Spark 1.2 and Muse Code beta today, with benchmark scores that put it near GPT-5.5 territory at a fraction of the cost — though the Contributor pricing tier's data-use terms deserve a close read before teams adopt it. Cloudflare dropped a coordinated set of releases rewriting how agents interact with the web: a stateless MCP spec, a V8-native headless browser, WebMCP for any hosted site, and an agent readiness scoring tool. Qwen3.8-Max open weights land next week, and a ChatGPT sandbox attack chain presented at Black Hat is worth understanding even though OpenAI patched the specific vector. DeepSeek signaled a significant price hike with no timeline, which matters given how many teams built cost assumptions around its rates.

---

## Meta releases Muse Spark 1.2 and Muse Code beta

Muse Spark 1.2 scores 54 on the Artificial Analysis Intelligence Index — roughly tied with GPT-5.5 and Grok 4.5, behind Claude Opus 5 (61) and GPT-5.6 Sol (59). Pricing holds at $1.25/$4.25 per million input/output tokens, putting it on the cost-efficiency frontier at $0.40 per task. The standout architectural detail in Muse Code is durable agent state: a local append-only event log records every model call, tool action, and approval, which is what let it sustain 1,000+ tool calls across a 24-hour GPU kernel optimization run. The Contributor pricing tier is 12–21x cheaper but funds Meta training on your prompts and completions — that's a governance decision, not just a billing one, and enterprise teams should make it deliberately rather than by default. [Read more](https://daily.dev/feed-by-ids?id=W7fg7hcmv&id=MLo83ZePz&id=zFef5liF7&id=Wb2YcAyN1&id=9I2Vnhl0S)

## Cloudflare rewrites the agent-web interface across five coordinated releases

Cloudflare shipped a lot today. MCP 2026-07-28 goes fully stateless — no sticky sessions, no open streams, each request carries its own context, which means MCP servers can run as ordinary HTTP workloads on Workers. Kitesurf is a new V8-isolate headless browser built for agents, not humans: 3–7x lower CPU and memory than Chromium, though ~1.7x slower wall time, already passing 215,000+ Web Platform Tests. WebMCP adds an MCP interface to any Cloudflare-hosted site via a single dashboard toggle, with Chrome 146 shipping the browser standard experimentally. AI Search now handles crawling, embedding, and retrieval automatically. And an Agent Readiness tool scores your site on how well AI agents can actually navigate it. The stateless MCP spec is the most consequential piece — Sentry and Linear are already running it in production. [Read more](https://daily.dev/feed-by-ids?id=zbwzrUpm5&id=PFNyPoUJE&id=h0QMMm3j8&id=cTSTzeRT3&id=o6iD7ejgp&id=Quu3Lagll)

## Qwen3.8-Max benchmarks and open weights timeline

Alibaba's Qwen3.8-Max is a 2.4T parameter MoE model (95B active per token) with a 1M token context window, priced at $2/M input and $6/M output. Benchmark highlights: 93% on PaperBench, 86.6% on Terminal-Bench 2.1, 67.7% on SWE-bench Pro. Open weights for both Qwen3.8-Max and a companion 27B checkpoint are scheduled for release the week of August 10 on Hugging Face — making it the first Max-class Qwen model with open weights. A 16-day unattended CLI run (265 commits, 127 PRs) and a 125-hour ML paper reproduction run are the more interesting capability demonstrations than the benchmark numbers alone. [Read more](https://daily.dev/feed-by-ids?id=qgszGYnxv&id=yBJ65jj6z&id=NayoeGGgd)

## ChatGPT sandbox attack chain demonstrated at Black Hat

Palo Alto Networks researcher Simcha Kosman presented a multi-step attack at Black Hat USA 2026 that achieved C2-style control over ChatGPT's isolated sandbox. The chain exploits a URL-based prompt execution flaw on iPhone/Mac, tricks ChatGPT into executing malicious spreadsheet code, stages data from connected Google Drive and Gmail, then uses a shared JFrog Artifactory backend as a covert channel — encoding binary data through account lockout states. OpenAI removed the vulnerable Artifactory behavior and addressed other findings within the 90-day disclosure window, and disputes that this constitutes a true sandbox escape. Worth understanding regardless: the attack surface for agents with connected tools is meaningfully larger than for standalone chat. [Read more](https://daily.dev/posts/Pb2mhJTIB)

---

## Also notable

- **DeepSeek signals significant API price hike with no timeline:** DeepSeek announced a forthcoming price increase for its API — no rates or date given — just days after releasing DeepSeek-V4-Flash-0731 at $0.14/M input tokens, which had become the fastest-growing model by token usage on Ollama; the likely cause is demand management rather than margin correction, but teams that built cost models around DeepSeek's rates should start stress-testing alternatives now. [Read more](https://daily.dev/feed-by-ids?id=PoU94OYEK&id=Sv0MjKIAd)
- **Grok 4.5 completed a one-shot task at $0.15 vs Kimi K3's $1.98 in head-to-head cost test:** An AI/ML API experiment running the same prompt across four models found Grok 4.5 at $0.15 per completed figure while Kimi K3 spent ~19 minutes thinking and billed $1.98 — 13x more — reinforcing that cost per completed task, not cost per token, is the metric that matters for production agent pipelines. [Read more](https://daily.dev/posts/gGmiOXo6y)
- **tl;dv left 181,874 meetings readable by any logged-in user for six months after disclosure:** Researcher BobDaHacker found tl;dv's Cloud Firestore had no tenant isolation on its meetings collection, exposing metadata for 181,874 meetings across 84,312 users and 35,003 domains — including government sessions from 23 countries — and was able to join active calls by impersonating an AI notetaker bot roughly 80% of the time; the CTO never responded and the vulnerability was still live at publication. [Read more](https://daily.dev/posts/RbI40l2h6)
- **Microsoft's open-source unit test agent hits 92.1% completion vs 78.9% for plain Copilot across 152 tasks:** Microsoft released code-testing-generator in the dotnet/skills repo, an agent that uses a research-plan-implement pipeline to write and validate unit tests across .NET, Python, Go, TypeScript, Java, and Rust — achieving 92.1% task completion and 63% fewer failures than stock GitHub Copilot, with most gains on vague or underspecified prompts. [Read more](https://daily.dev/posts/QnRzDAIsK)
- **Nvidia paper: cross-model KV cache transfer runs 2.7–25x faster with 73–98% accuracy retention on 4 of 6 model pairs:** Nvidia researchers show a closed-form linear mapper trained on 500 calibration sequences (no backpropagation) can transfer KV cache between related models in the Qwen3, Llama 3.1, and Ministral families, eliminating full prompt reprocessing when switching models mid-session — with accuracy loss depending on whether errors land in attention-critical positions, not on total conversion error. [Read more](https://daily.dev/posts/emY4sIPR4)

## Similar posts on daily.dev

- [Meta enters the crowded AI coding battle with Muse Spark 1.1](https://daily.dev/posts/meta-enters-the-crowded-ai-coding-battle-with-muse-spark-1-1-wtfjrleyt) · TechCrunch · 1 upvotes · 0 comments

---

Tags: [#security](https://daily.dev/tags/security), [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#mcp](https://daily.dev/tags/mcp), [#cloudflare](https://daily.dev/tags/cloudflare)

[View this post on daily.dev](https://daily.dev/posts/meta-muse-spark-1-2-benchmarks-cloudflare-rewrites-the-agentic-web-stack-gabz2pamn)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"DiscussionForumPosting","mainEntityOfPage":"https://daily.dev/posts/meta-muse-spark-1-2-benchmarks-cloudflare-rewrites-the-agentic-web-stack-gabz2pamn","headline":"Meta Muse Spark 1.2 benchmarks, Cloudflare rewrites the agentic web stack","text":"Meta released Muse Spark 1.2 and Muse Code beta today, with benchmark scores that put it near GPT-5.5 territory at a fraction of the cost — though the Contributor pricing tier's data-use terms deserve a close read before teams adopt it. Cloudflare dropped a coordinated set of releases rewriting how agents interact with the web: a stateless MCP spec, a V8-native headless browser, WebMCP for any hosted site, and an agent readiness scoring tool. Qwen3.8-Max open weights land next week, and a ChatGPT sandbox attack chain presented at Black Hat is worth understanding even though OpenAI patched the specific vector. DeepSeek signaled a significant price hike with no timeline, which matters given how many teams built cost assumptions around its rates.","url":"https://daily.dev/posts/meta-muse-spark-1-2-benchmarks-cloudflare-rewrites-the-agentic-web-stack-gabz2pamn","datePublished":"2026-08-07T04:18:14.649Z","dateModified":"2026-08-07T04:18:54.070Z","author":{"@type":"Organization","name":"Agentic Digest","logo":"https://media.daily.dev/image/upload/s--V91DY4ls--/f_auto,q_auto/v1772617267/logos/agents_digest","url":"https://daily.dev/sources/agents_digest"},"interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":3},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"isPartOf":{"@type":"WebPage","url":"https://daily.dev/sources/agents_digest","name":"Agentic Digest"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Agentic Digest","item":"https://daily.dev/sources/agents_digest"},{"@type":"ListItem","position":3,"name":"Meta Muse Spark 1.2 benchmarks, Cloudflare rewrites the agentic web stack"}]}
```

