<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/qwen3-8-flash-next-mtplx-2-10-insanely-fast-local-coding-agents--vmnhnyoya" -->

---
title: Qwen3.8-Flash-Next + MTPLX 2.10: INSANELY Fast Local...
description: A creator demonstrates MLX 2.10 (referred to as MTPLX) running the Qwen3.8-Flash-Next model locally on an Apple M5 Max with 128GB RAM, hitting up to 70...
canonical: https://daily.dev/posts/qwen3-8-flash-next-mtplx-2-10-insanely-fast-local-coding-agents--vmnhnyoya
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Qwen3.8-Flash-Next + MTPLX 2.10: INSANELY Fast Local Coding Agents! | daily.dev
og:description: A creator demonstrates MLX 2.10 (referred to as MTPLX) running the Qwen3.8-Flash-Next model locally on an Apple M5 Max with 128GB RAM, hitting up to 70...
og:url: https://daily.dev/posts/qwen3-8-flash-next-mtplx-2-10-insanely-fast-local-coding-agents--vmnhnyoya
og:image: https://api.daily.dev/og/posts/vmNhnyoyA.png
og:image:alt: Qwen3.8-Flash-Next + MTPLX 2.10: INSANELY Fast Local Coding Agents!
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Qwen3.8-Flash-Next + MTPLX 2.10: INSANELY Fast Local Coding Agents!

**[Execute Automation](https://daily.dev/sources/executeautomation)** · 14 min read · 3 upvotes · 0 comments

## Summary

A creator demonstrates MLX 2.10 (referred to as MTPLX) running the Qwen3.8-Flash-Next model locally on an Apple M5 Max with 128GB RAM, hitting up to 70 tokens/sec. The video walks through five new features in this release: Qwen3.8-Flash-Next support with draft tokens/MTP, NVMe SSD offloading with N-gram support, long context handling past 147K tokens with zero collapse, memory-aware KV caching with SSD spill, and faster agent loop support for coding agents. The creator compares 4-bit vs 8-bit quantized models, runs a .NET 8 to .NET 10 migration test with a coding agent (completed in 5 minutes), and compares image generation quality and speed against the older MLX version, finding MLX 2.10 much faster but with somewhat lower image quality.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=jFqrXIjttE8>

## Questions this post answers

### What new features does MLX 2.10 add for running Qwen3.8-Flash-Next locally on Apple Silicon?

MLX version 2.10 adds five major features: Qwen3.8-Flash-Next support with draft token and MTP acceleration, NVMe SSD offloading with N-gram support so the full model doesn't need to fit in RAM, long context optimization supporting over 147,000 tokens without collapse, memory-aware KV caching with controlled SSD spill to avoid swap slowdowns, and improved agent loop support for faster coding agent inferencing and tool calling.

_Developers tuning local inference setups can track framework updates like this one on daily.dev._

### How fast is Qwen3.8-Flash-Next running locally on an Apple M5 Max with 128GB RAM using MLX 2.10?

Qwen3.8-Flash-Next reaches around 60-73 tokens per second on an Apple M5 Max with 128GB RAM using MLX 2.10's 4-bit quantized speed model, and 50-60 tokens per second when used inside a coding agent. A .NET 8 to .NET 10 migration task completed in about 5 minutes using this setup with a coding agent.

_Anyone benchmarking local coding agent throughput can follow real-world numbers like these on daily.dev._

### What is the tradeoff between MLX and the older MLX version when running the same quantized model for coding versus image generation?

MLX 2.10 delivers much faster inference for coding agent tasks, completing a migration in roughly 5 minutes versus long delays with the older version, but produces noticeably lower quality image generation output. The older MLX version took about 21 minutes to generate a test image compared to roughly 1-2 minutes with MLX 2.10, but yielded richer, higher quality images for the same 4-bit quantized Qwen3.8-Flash-Next model.

_Weighing speed against output quality across local inference tools is easier when tracking comparisons on daily.dev._

## Similar posts on daily.dev

- [qMLX: Maximising my AI psychosis by minmaxing my Mac Studio](https://daily.dev/posts/qmlx-maximising-my-ai-psychosis-by-minmaxing-my-mac-studio-gigqbqete) · Hacker News · 1 upvotes · 0 comments

---

Tags: [#data-science](https://daily.dev/tags/data-science), [#ai-agents](https://daily.dev/tags/ai-agents), [#qwen](https://daily.dev/tags/qwen)

[View this post on daily.dev](https://daily.dev/posts/qwen3-8-flash-next-mtplx-2-10-insanely-fast-local-coding-agents--vmnhnyoya)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Qwen3.8-Flash-Next + MTPLX 2.10: INSANELY Fast Local Coding Agents!","url":"https://daily.dev/posts/qwen3-8-flash-next-mtplx-2-10-insanely-fast-local-coding-agents--vmnhnyoya","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/qwen3-8-flash-next-mtplx-2-10-insanely-fast-local-coding-agents--vmnhnyoya"},"datePublished":"2026-09-01T14:22:58.205Z","dateModified":"2026-09-01T14:23:23.970Z","description":"A creator demonstrates MLX 2.10 (referred to as MTPLX) running the Qwen3.8-Flash-Next model locally on an Apple M5 Max with 128GB RAM, hitting up to 70...","image":"https://i.ytimg.com/vi/jFqrXIjttE8/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/jFqrXIjttE8/sddefault.jpg","isAccessibleForFree":true,"articleSection":"Execute Automation","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Execute Automation","logo":"https://media.daily.dev/image/upload/s--Uh-EE9Cf--/f_auto,q_auto/v1780213688/logos/executeautomation?_a=BAMAMiWQ0","url":"https://daily.dev/sources/executeautomation"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/qwen3-8-flash-next-mtplx-2-10-insanely-fast-local-coding-agents--vmnhnyoya","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":3},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"data-science,ai-agents,qwen","timeRequired":"PT14M","video":{"@type":"VideoObject","name":"Qwen3.8-Flash-Next + MTPLX 2.10: INSANELY Fast Local Coding Agents!","description":"A creator demonstrates MLX 2.10 (referred to as MTPLX) running the Qwen3.8-Flash-Next model locally on an Apple M5 Max with 128GB RAM, hitting up to 70...","thumbnailUrl":"https://i.ytimg.com/vi/jFqrXIjttE8/sddefault.jpg","uploadDate":"2026-09-01T14:22:58.205Z","duration":"PT14M","url":"https://api.daily.dev/r/vmNhnyoyA","embedUrl":"https://www.youtube.com/embed/jFqrXIjttE8"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Execute Automation","item":"https://daily.dev/sources/executeautomation"},{"@type":"ListItem","position":3,"name":"Qwen3.8-Flash-Next + MTPLX 2.10: INSANELY Fast Local Coding Agents!"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/qwen3-8-flash-next-mtplx-2-10-insanely-fast-local-coding-agents--vmnhnyoya#faq","mainEntity":[{"@type":"Question","name":"What new features does MLX 2.10 add for running Qwen3.8-Flash-Next locally on Apple Silicon?","acceptedAnswer":{"@type":"Answer","text":"MLX version 2.10 adds five major features: Qwen3.8-Flash-Next support with draft token and MTP acceleration, NVMe SSD offloading with N-gram support so the full model doesn't need to fit in RAM, long context optimization supporting over 147,000 tokens without collapse, memory-aware KV caching with controlled SSD spill to avoid swap slowdowns, and improved agent loop support for faster coding agent inferencing and tool calling. Developers tuning local inference setups can track framework updates like this one on daily.dev."}},{"@type":"Question","name":"How fast is Qwen3.8-Flash-Next running locally on an Apple M5 Max with 128GB RAM using MLX 2.10?","acceptedAnswer":{"@type":"Answer","text":"Qwen3.8-Flash-Next reaches around 60-73 tokens per second on an Apple M5 Max with 128GB RAM using MLX 2.10's 4-bit quantized speed model, and 50-60 tokens per second when used inside a coding agent. A .NET 8 to .NET 10 migration task completed in about 5 minutes using this setup with a coding agent. Anyone benchmarking local coding agent throughput can follow real-world numbers like these on daily.dev."}},{"@type":"Question","name":"What is the tradeoff between MLX and the older MLX version when running the same quantized model for coding versus image generation?","acceptedAnswer":{"@type":"Answer","text":"MLX 2.10 delivers much faster inference for coding agent tasks, completing a migration in roughly 5 minutes versus long delays with the older version, but produces noticeably lower quality image generation output. The older MLX version took about 21 minutes to generate a test image compared to roughly 1-2 minutes with MLX 2.10, but yielded richer, higher quality images for the same 4-bit quantized Qwen3.8-Flash-Next model. Weighing speed against output quality across local inference tools is easier when tracking comparisons on daily.dev."}}]}
```

