<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/qwen3-8-27b-low-vs-xhigh-thinking-dflash2-speculative-decoding-on-m5-max--ud8lzf0bl" -->

---
title: Qwen3.8-27B: Low vs xHigh Thinking + DFlash2 Speculative...
description: A hands-on comparison testing Qwen3.8-27B&#x27;s thinking modes (low, medium, high, xhigh) alongside DFlash2 speculative decoding on an M5 Max MacBook Pro. The...
canonical: https://daily.dev/posts/qwen3-8-27b-low-vs-xhigh-thinking-dflash2-speculative-decoding-on-m5-max--ud8lzf0bl
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Qwen3.8-27B: Low vs xHigh Thinking + DFlash2 Speculative Decoding on M5 Max ⚡️ | daily.dev
og:description: A hands-on comparison testing Qwen3.8-27B&#x27;s thinking modes (low, medium, high, xhigh) alongside DFlash2 speculative decoding on an M5 Max MacBook Pro. The...
og:url: https://daily.dev/posts/qwen3-8-27b-low-vs-xhigh-thinking-dflash2-speculative-decoding-on-m5-max--ud8lzf0bl
og:image: https://api.daily.dev/og/posts/ud8lZF0BL.png
og:image:alt: Qwen3.8-27B: Low vs xHigh Thinking + DFlash2 Speculative Decoding on M5 Max ⚡️
og:image:width: 1200
og:image:height: 630
og:locale: en
---

[Execute Automation](https://daily.dev/sources/executeautomation)

[Watch video](https://api.daily.dev/r/ud8lZF0BL)

# [Qwen3.8-27B: Low vs xHigh Thinking + DFlash2 Speculative Decoding on M5 Max ⚡️](https://api.daily.dev/r/ud8lZF0BL "Go to post")

A hands-on comparison testing Qwen3.8-27B's thinking modes (low, medium, high, xhigh) alongside DFlash2 speculative decoding on an M5 Max MacBook Pro. The DFlash2 setup in llama.cpp (via an unmerged pull request, GGUF format only) is benchmarked against a plain Ollama MLX setup across different thinking levels, using a coding agent. Despite DFlash2's advertised 70 tokens/sec claim, actual results ranged from 11-35 tokens/sec depending on thinking level, while Ollama surprisingly outperformed it, hitting up to 56 tokens/sec on low settings and 24-47 tokens/sec on xhigh. The creator concludes medium thinking level offers a good balance, with xhigh giving noticeably richer output at the cost of speed.

[#ai-inference](/tags/ai-inference "Check all #ai-inference posts")[#ollama](/tags/ollama "Check all #ollama posts")[#qwen](/tags/qwen "Check all #qwen posts")[#llama-cpp](/tags/llama-cpp "Check all #llama-cpp posts")

Aug 20 • 16m watch time

Questions this post answers

Does DFlash2 speculative decoding actually give faster token generation than Ollama for Qwen3.8-27B on Apple Silicon?

Not necessarily. In direct testing on an M5 Max MacBook Pro, DFlash2 speculative decoding in llama.cpp produced only 11-35 tokens per second across different thinking levels, far below the advertised 70 tokens per second, while a plain Ollama MLX setup without speculative decoding reached 24-56 tokens per second, outperforming the draft-token approach in most tests. Anyone weighing local LLM inference setups can compare real-world speculative decoding results on daily.dev.

Is DFlash2 speculative decoding merged into llama.cpp yet?

No, DFlash2 support is not yet merged into llama.cpp mainline; it exists only as an open pull request. To use it, developers need to clone the llama.cpp repository, check out that pull request's branch, and build it themselves, and it only works with GGUF format models, not MLX. Track llama.cpp feature merges like this on daily.dev before betting a workflow on unmerged code.

What thinking level should I use for Qwen3.8-27B to balance speed and output quality?

Medium thinking level offers a good balance between speed and detail for Qwen3.8-27B. Low settings produce faster output but with less detail and analysis, while xhigh produces significantly richer, more thorough responses at the cost of dropping to as low as 11-24 tokens per second depending on the backend used. Developers tuning coding agent thinking levels can dig into benchmarks like this on daily.dev.

Comment

Bookmark

Copy

![Placeholder image for anonymous user](https://media.daily.dev/image/upload/s--qsFuKGv_--/t_logo,f_auto/public/noProfile)Share your thoughts Post

Share this post

[![Execute Automation's image](https://media.daily.dev/image/upload/s--Uh-EE9Cf--/f_auto,q_auto/v1780213688/logos/executeautomation?_a=BAMAMiWQ0)](https://daily.dev/sources/executeautomation)

[Execute Automation](https://daily.dev/sources/executeautomation "https://daily.dev/sources/executeautomation")

5 Followers

•

27 Upvotes

#### Would you recommend this post?

Copy link

Slack

WhatsApp

Facebook

X

New Squad

Copy linkShare with your friends

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Qwen3.8-27B: Low vs xHigh Thinking + DFlash2 Speculative Decoding on M5 Max ⚡️","url":"https://daily.dev/posts/qwen3-8-27b-low-vs-xhigh-thinking-dflash2-speculative-decoding-on-m5-max--ud8lzf0bl","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/qwen3-8-27b-low-vs-xhigh-thinking-dflash2-speculative-decoding-on-m5-max--ud8lzf0bl"},"datePublished":"2026-08-20T14:26:55.907Z","dateModified":"2026-08-20T14:27:52.442Z","description":"A hands-on comparison testing Qwen3.8-27B's thinking modes (low, medium, high, xhigh) alongside DFlash2 speculative decoding on an M5 Max MacBook Pro. The...","image":"https://i.ytimg.com/vi/nz8WASzfSPs/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/nz8WASzfSPs/sddefault.jpg","isAccessibleForFree":true,"articleSection":"Execute Automation","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Execute Automation","logo":"https://media.daily.dev/image/upload/s--Uh-EE9Cf--/f_auto,q_auto/v1780213688/logos/executeautomation?_a=BAMAMiWQ0","url":"https://daily.dev/sources/executeautomation"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/qwen3-8-27b-low-vs-xhigh-thinking-dflash2-speculative-decoding-on-m5-max--ud8lzf0bl","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-inference,ollama,qwen,llama-cpp","timeRequired":"PT16M","video":{"@type":"VideoObject","name":"Qwen3.8-27B: Low vs xHigh Thinking + DFlash2 Speculative Decoding on M5 Max ⚡️","description":"A hands-on comparison testing Qwen3.8-27B's thinking modes (low, medium, high, xhigh) alongside DFlash2 speculative decoding on an M5 Max MacBook Pro. The...","thumbnailUrl":"https://i.ytimg.com/vi/nz8WASzfSPs/sddefault.jpg","uploadDate":"2026-08-20T14:26:55.907Z","duration":"PT16M","url":"https://api.daily.dev/r/ud8lZF0BL","embedUrl":"https://www.youtube.com/embed/nz8WASzfSPs"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Execute Automation","item":"https://daily.dev/sources/executeautomation"},{"@type":"ListItem","position":3,"name":"Qwen3.8-27B: Low vs xHigh Thinking + DFlash2 Speculative Decoding on M5 Max ⚡️"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/qwen3-8-27b-low-vs-xhigh-thinking-dflash2-speculative-decoding-on-m5-max--ud8lzf0bl#faq","mainEntity":[{"@type":"Question","name":"Does DFlash2 speculative decoding actually give faster token generation than Ollama for Qwen3.8-27B on Apple Silicon?","acceptedAnswer":{"@type":"Answer","text":"Not necessarily. In direct testing on an M5 Max MacBook Pro, DFlash2 speculative decoding in llama.cpp produced only 11-35 tokens per second across different thinking levels, far below the advertised 70 tokens per second, while a plain Ollama MLX setup without speculative decoding reached 24-56 tokens per second, outperforming the draft-token approach in most tests. Anyone weighing local LLM inference setups can compare real-world speculative decoding results on daily.dev."}},{"@type":"Question","name":"Is DFlash2 speculative decoding merged into llama.cpp yet?","acceptedAnswer":{"@type":"Answer","text":"No, DFlash2 support is not yet merged into llama.cpp mainline; it exists only as an open pull request. To use it, developers need to clone the llama.cpp repository, check out that pull request's branch, and build it themselves, and it only works with GGUF format models, not MLX. Track llama.cpp feature merges like this on daily.dev before betting a workflow on unmerged code."}},{"@type":"Question","name":"What thinking level should I use for Qwen3.8-27B to balance speed and output quality?","acceptedAnswer":{"@type":"Answer","text":"Medium thinking level offers a good balance between speed and detail for Qwen3.8-27B. Low settings produce faster output but with less detail and analysis, while xhigh produces significantly richer, more thorough responses at the cost of dropping to as low as 11-24 tokens per second depending on the backend used. Developers tuning coding agent thinking levels can dig into benchmarks like this on daily.dev."}}]}
```

