<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/muse-spark-1-3-gemini-3-8-flash-gemini-has-leveled-up-big-time--s93qzvefh" -->

---
title: Muse Spark 1.3 &amp; Gemini 3.8 Flash: Gemini has leveled up...
description: A YouTube-style benchmark review compares two new AI model releases, Meta&#x27;s Muse Spark 1.3 and Google&#x27;s Gemini 3.8 Flash, using a custom 8-question...
canonical: https://daily.dev/posts/muse-spark-1-3-gemini-3-8-flash-gemini-has-leveled-up-big-time--s93qzvefh
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Muse Spark 1.3 &amp; Gemini 3.8 Flash: Gemini has leveled up BIG TIME! | daily.dev
og:description: A YouTube-style benchmark review compares two new AI model releases, Meta&#x27;s Muse Spark 1.3 and Google&#x27;s Gemini 3.8 Flash, using a custom 8-question...
og:url: https://daily.dev/posts/muse-spark-1-3-gemini-3-8-flash-gemini-has-leveled-up-big-time--s93qzvefh
og:image: https://api.daily.dev/og/posts/s93qZvefh.png
og:image:alt: Muse Spark 1.3 &amp; Gemini 3.8 Flash: Gemini has leveled up BIG TIME!
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Muse Spark 1.3 & Gemini 3.8 Flash: Gemini has leveled up BIG TIME!

**[AICodeKing](https://daily.dev/sources/aicodeking)** · 10 min read · 0 upvotes · 0 comments

## Summary

A YouTube-style benchmark review compares two new AI model releases, Meta's Muse Spark 1.3 and Google's Gemini 3.8 Flash, using a custom 8-question coding/agentic benchmark called Kingbench 3. Gemini 3.8 Flash jumps dramatically from its predecessor (30% to 81.25%), landing in the top five of the leaderboard, while Muse Spark 1.3 actually regresses from 1.2's 76.25% down to 71.25%, losing visual/front-end strength despite gains on logic tasks and a record wristwatch score. Both models still exhibit a file-overwriting problem when used for coding edits. The reviewer recommends Gemini 3.8 Flash via Google's Antigravity tool for free research/coding work, and mentions both models are available on a third-party agentic coding workspace called Verdant.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=WZDtEAFHj7k>

## Questions this post answers

### How does Gemini 3.8 Flash compare to Gemini 3.5 Flash on coding benchmarks?

Gemini 3.8 Flash scores 81.25% (65/80) on the Kingbench 3 benchmark, compared to just 30% for Gemini 3.5 Flash, a jump of over 51 percentage points in one generation. That score places it in the top five of the leaderboard, tied with Qwen 3.8 Max, just below Fable 5 (82.5%) and above Opus 4.8 (80%).

_Developers tracking flash-tier model upgrades can follow model benchmark shifts like this on daily.dev._

### Is Muse Spark 1.3 better or worse than Muse Spark 1.2 for coding tasks?

Muse Spark 1.3 is worse overall, scoring 71.25% (57/80) on Kingbench 3 versus 76.25% for version 1.2, a five-point regression. It improved on the elevator simulation and set a new record on the 3D wristwatch test, but fell off hard on the SVG panda test (10 down to 5) and the folding table test, the areas that made 1.2 stand out.

_Teams deciding whether to upgrade to Muse Spark 1.3 can weigh benchmark tradeoffs like these on daily.dev._

### Do Gemini 3.8 Flash and Muse Spark 1.3 have issues with overwriting files during coding tasks?

Yes, both models tend to overwrite entire files instead of making targeted edits when asked for small changes. This overwriting habit was already present in Muse Spark 1.2 and remains unfixed in 1.3, while Gemini 3.8 Flash has newly picked up the same behavior, unlike earlier Gemini models. Developers are advised to keep commits small and review diffs carefully.

_Developers wary of AI agents nuking unrelated code changes can track model quirks like this on daily.dev._

## Similar posts on daily.dev

- [Google’s Gemini 3.5 Flash beats the frontier models](https://daily.dev/posts/google-s-gemini-3-5-flash-beats-the-frontier-models-gdirdibym) · The New Stack · 0 upvotes · 0 comments
- [Google releases Gemini 3 Flash, enabling faster, more cost effective reasoning](https://daily.dev/posts/google-releases-gemini-3-flash-enabling-faster-more-cost-effective-reasoning-1y9yu2yfb) · SD Times · 0 upvotes · 0 comments

---

Tags: [#ai](https://daily.dev/tags/ai), [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#google-gemini](https://daily.dev/tags/google-gemini)

[View this post on daily.dev](https://daily.dev/posts/muse-spark-1-3-gemini-3-8-flash-gemini-has-leveled-up-big-time--s93qzvefh)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Muse Spark 1.3 & Gemini 3.8 Flash: Gemini has leveled up BIG TIME!","url":"https://daily.dev/posts/muse-spark-1-3-gemini-3-8-flash-gemini-has-leveled-up-big-time--s93qzvefh","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/muse-spark-1-3-gemini-3-8-flash-gemini-has-leveled-up-big-time--s93qzvefh"},"datePublished":"2026-09-03T10:26:42.617Z","dateModified":"2026-09-03T10:27:43.022Z","description":"A YouTube-style benchmark review compares two new AI model releases, Meta's Muse Spark 1.3 and Google's Gemini 3.8 Flash, using a custom 8-question...","image":"https://i.ytimg.com/vi/WZDtEAFHj7k/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/WZDtEAFHj7k/sddefault.jpg","isAccessibleForFree":true,"articleSection":"AICodeKing","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"AICodeKing","logo":"https://media.daily.dev/image/upload/s--x7nDUfWj--/f_auto,q_auto/v1768208312/logos/aicodeking","url":"https://daily.dev/sources/aicodeking"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/muse-spark-1-3-gemini-3-8-flash-gemini-has-leveled-up-big-time--s93qzvefh","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,llm,ai-agents,google-gemini","timeRequired":"PT10M","video":{"@type":"VideoObject","name":"Muse Spark 1.3 & Gemini 3.8 Flash: Gemini has leveled up BIG TIME!","description":"A YouTube-style benchmark review compares two new AI model releases, Meta's Muse Spark 1.3 and Google's Gemini 3.8 Flash, using a custom 8-question...","thumbnailUrl":"https://i.ytimg.com/vi/WZDtEAFHj7k/sddefault.jpg","uploadDate":"2026-09-03T10:26:42.617Z","duration":"PT10M","url":"https://api.daily.dev/r/s93qZvefh","embedUrl":"https://www.youtube.com/embed/WZDtEAFHj7k"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"AICodeKing","item":"https://daily.dev/sources/aicodeking"},{"@type":"ListItem","position":3,"name":"Muse Spark 1.3 & Gemini 3.8 Flash: Gemini has leveled up BIG TIME!"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/muse-spark-1-3-gemini-3-8-flash-gemini-has-leveled-up-big-time--s93qzvefh#faq","mainEntity":[{"@type":"Question","name":"How does Gemini 3.8 Flash compare to Gemini 3.5 Flash on coding benchmarks?","acceptedAnswer":{"@type":"Answer","text":"Gemini 3.8 Flash scores 81.25% (65/80) on the Kingbench 3 benchmark, compared to just 30% for Gemini 3.5 Flash, a jump of over 51 percentage points in one generation. That score places it in the top five of the leaderboard, tied with Qwen 3.8 Max, just below Fable 5 (82.5%) and above Opus 4.8 (80%). Developers tracking flash-tier model upgrades can follow model benchmark shifts like this on daily.dev."}},{"@type":"Question","name":"Is Muse Spark 1.3 better or worse than Muse Spark 1.2 for coding tasks?","acceptedAnswer":{"@type":"Answer","text":"Muse Spark 1.3 is worse overall, scoring 71.25% (57/80) on Kingbench 3 versus 76.25% for version 1.2, a five-point regression. It improved on the elevator simulation and set a new record on the 3D wristwatch test, but fell off hard on the SVG panda test (10 down to 5) and the folding table test, the areas that made 1.2 stand out. Teams deciding whether to upgrade to Muse Spark 1.3 can weigh benchmark tradeoffs like these on daily.dev."}},{"@type":"Question","name":"Do Gemini 3.8 Flash and Muse Spark 1.3 have issues with overwriting files during coding tasks?","acceptedAnswer":{"@type":"Answer","text":"Yes, both models tend to overwrite entire files instead of making targeted edits when asked for small changes. This overwriting habit was already present in Muse Spark 1.2 and remains unfixed in 1.3, while Gemini 3.8 Flash has newly picked up the same behavior, unlike earlier Gemini models. Developers are advised to keep commits small and review diffs carefully. Developers wary of AI agents nuking unrelated code changes can track model quirks like this on daily.dev."}}]}
```

