<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/mac-mini-m6-first-look-qwen3-8-27b-hits-53-tps-locally-eqglitifc" -->

---
title: Mac mini M6 First Look — Qwen3.8-27B Hits 53+ TPS Locally
description: A first-look review of the new Mac Mini M6, just launched globally, testing local LLM inference performance. The reviewer runs the Qwen3.8-27B dense model...
canonical: https://daily.dev/posts/mac-mini-m6-first-look-qwen3-8-27b-hits-53-tps-locally-eqglitifc
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Mac mini M6 First Look — Qwen3.8-27B Hits 53+ TPS Locally | daily.dev
og:description: A first-look review of the new Mac Mini M6, just launched globally, testing local LLM inference performance. The reviewer runs the Qwen3.8-27B dense model...
og:url: https://daily.dev/posts/mac-mini-m6-first-look-qwen3-8-27b-hits-53-tps-locally-eqglitifc
og:image: https://api.daily.dev/og/posts/EQGLITIfC.png
og:image:alt: Mac mini M6 First Look — Qwen3.8-27B Hits 53+ TPS Locally
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Mac mini M6 First Look — Qwen3.8-27B Hits 53+ TPS Locally

**[Execute Automation](https://daily.dev/sources/executeautomation)** · 12 min read · 0 upvotes · 0 comments

## Summary

A first-look review of the new Mac Mini M6, just launched globally, testing local LLM inference performance. The reviewer runs the Qwen3.8-27B dense model using LM Studio with different engines (MLX and a separate 'splash' engine) and quantization settings, showing token-per-second speeds ranging from about 8-9 tps with plain MLX, up to 17-19 tps with MTP (multi-token prediction) enabled, and 26-56 tps using the splash engine. The M6 features 12 CPU cores, 12 GPU cores, dual neural engines, and up to 170 Gbps memory bandwidth (up from 120 Gbps on M4), with configurations now capped at 32GB instead of the M4's prior 32GB tier starting point being replaced by a 24GB base. The reviewer notes the M4 Mac Mini could not even load this 27B model, calling the M6 a major leap for local AI inferencing, though price increased by about $200.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=7ugFOo8e8H4>

## Questions this post answers

### How many tokens per second does the Mac Mini M6 get running the Qwen3.8-27B model locally?

Performance varies significantly by engine and setting. With plain MLX in LM Studio, the Mac Mini M6 (32GB RAM) gets roughly 8-9 tokens per second on the Qwen3.8-27B dense 4-bit quantized model; enabling lightning MTP (multi-token prediction) raises this to about 16-19 tokens per second; using a separate splash engine instead of MLX pushes speeds up to 26-56 tokens per second depending on the prompt.

_Developers benchmarking local LLM hardware choices can track real-world Apple Silicon performance reports on daily.dev._

### What are the hardware differences between the Apple M6 Mac Mini and the M4 Mac Mini?

The M6 Mac Mini adds a dual neural engine for faster AI processing, 12 CPU cores, 12 GPU cores, and up to 170 Gbps of memory bandwidth compared to 120 Gbps on the M4. RAM configuration also shifted: the M4 previously offered a 32GB tier, but the base M6 now starts at 24GB with a 32GB option available. The price increased by about $200 due to the memory upgrade.

_Anyone comparing Apple Silicon chips before buying can follow detailed hardware comparisons on daily.dev._

### Can the Mac Mini M4 run the Qwen3.8-27B parameter dense model locally?

No, the Mac Mini M4 could not even load the Qwen3.8-27B parameter dense model, making it unusable for this workload. Apple has stated the M6 delivers up to four times faster performance than the M4 for such AI inferencing tasks, which is reflected in the M6 successfully loading and running the model at usable speeds.

_Developers deciding which local hardware can handle large dense models can compare results like this on daily.dev._

---

Tags: [#lm-studio](https://daily.dev/tags/lm-studio)

[View this post on daily.dev](https://daily.dev/posts/mac-mini-m6-first-look-qwen3-8-27b-hits-53-tps-locally-eqglitifc)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Mac mini M6 First Look — Qwen3.8-27B Hits 53+ TPS Locally","url":"https://daily.dev/posts/mac-mini-m6-first-look-qwen3-8-27b-hits-53-tps-locally-eqglitifc","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/mac-mini-m6-first-look-qwen3-8-27b-hits-53-tps-locally-eqglitifc"},"datePublished":"2026-09-22T13:27:36.770Z","dateModified":"2026-09-22T13:28:00.904Z","description":"A first-look review of the new Mac Mini M6, just launched globally, testing local LLM inference performance. The reviewer runs the Qwen3.8-27B dense model...","image":"https://i.ytimg.com/vi/7ugFOo8e8H4/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/7ugFOo8e8H4/sddefault.jpg","isAccessibleForFree":true,"articleSection":"Execute Automation","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Execute Automation","logo":"https://media.daily.dev/image/upload/s--Uh-EE9Cf--/f_auto,q_auto/v1780213688/logos/executeautomation?_a=BAMAMiWQ0","url":"https://daily.dev/sources/executeautomation"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/mac-mini-m6-first-look-qwen3-8-27b-hits-53-tps-locally-eqglitifc","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"lm-studio","timeRequired":"PT12M","video":{"@type":"VideoObject","name":"Mac mini M6 First Look — Qwen3.8-27B Hits 53+ TPS Locally","description":"A first-look review of the new Mac Mini M6, just launched globally, testing local LLM inference performance. The reviewer runs the Qwen3.8-27B dense model...","thumbnailUrl":"https://i.ytimg.com/vi/7ugFOo8e8H4/sddefault.jpg","uploadDate":"2026-09-22T13:27:36.770Z","duration":"PT12M","url":"https://api.daily.dev/r/EQGLITIfC","embedUrl":"https://www.youtube.com/embed/7ugFOo8e8H4"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Execute Automation","item":"https://daily.dev/sources/executeautomation"},{"@type":"ListItem","position":3,"name":"Mac mini M6 First Look — Qwen3.8-27B Hits 53+ TPS Locally"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/mac-mini-m6-first-look-qwen3-8-27b-hits-53-tps-locally-eqglitifc#faq","mainEntity":[{"@type":"Question","name":"How many tokens per second does the Mac Mini M6 get running the Qwen3.8-27B model locally?","acceptedAnswer":{"@type":"Answer","text":"Performance varies significantly by engine and setting. With plain MLX in LM Studio, the Mac Mini M6 (32GB RAM) gets roughly 8-9 tokens per second on the Qwen3.8-27B dense 4-bit quantized model; enabling lightning MTP (multi-token prediction) raises this to about 16-19 tokens per second; using a separate splash engine instead of MLX pushes speeds up to 26-56 tokens per second depending on the prompt. Developers benchmarking local LLM hardware choices can track real-world Apple Silicon performance reports on daily.dev."}},{"@type":"Question","name":"What are the hardware differences between the Apple M6 Mac Mini and the M4 Mac Mini?","acceptedAnswer":{"@type":"Answer","text":"The M6 Mac Mini adds a dual neural engine for faster AI processing, 12 CPU cores, 12 GPU cores, and up to 170 Gbps of memory bandwidth compared to 120 Gbps on the M4. RAM configuration also shifted: the M4 previously offered a 32GB tier, but the base M6 now starts at 24GB with a 32GB option available. The price increased by about $200 due to the memory upgrade. Anyone comparing Apple Silicon chips before buying can follow detailed hardware comparisons on daily.dev."}},{"@type":"Question","name":"Can the Mac Mini M4 run the Qwen3.8-27B parameter dense model locally?","acceptedAnswer":{"@type":"Answer","text":"No, the Mac Mini M4 could not even load the Qwen3.8-27B parameter dense model, making it unusable for this workload. Apple has stated the M6 delivers up to four times faster performance than the M4 for such AI inferencing tasks, which is reflected in the M6 successfully loading and running the model at usable speeds. Developers deciding which local hardware can handle large dense models can compare results like this on daily.dev."}}]}
```

