<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/karpathy-s-trick-for-faster-local-llms-finally-has-proper-tooling-heblac4j0" -->

---
title: Karpathy’s Trick for Faster Local LLMs Finally Has...
description: On-device LLM inference at batch size one is bottlenecked by memory bandwidth rather than compute, since every decoded token requires streaming the full model...
canonical: https://daily.dev/posts/karpathy-s-trick-for-faster-local-llms-finally-has-proper-tooling-heblac4j0
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Karpathy’s Trick for Faster Local LLMs Finally Has Proper Tooling | daily.dev
og:description: On-device LLM inference at batch size one is bottlenecked by memory bandwidth rather than compute, since every decoded token requires streaming the full model...
og:url: https://daily.dev/posts/karpathy-s-trick-for-faster-local-llms-finally-has-proper-tooling-heblac4j0
og:image: https://api.daily.dev/og/posts/HEBLAc4j0.png
og:image:alt: Karpathy’s Trick for Faster Local LLMs Finally Has Proper Tooling
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Karpathy’s Trick for Faster Local LLMs Finally Has Proper Tooling

**[Daily Dose of Data Science \| Avi Chawla \| Substack](https://daily.dev/sources/dailydoseofds)** · 20 min read · 1 upvotes · 0 comments

## Summary

On-device LLM inference at batch size one is bottlenecked by memory bandwidth rather than compute, since every decoded token requires streaming the full model through memory. Uzu, an open-source local inference engine, targets this by combining 4-bit/8-bit integer quantization with a Hadamard transform to reduce rounding error, and a speculative decoding scheme (DFlash plus the Weaver adapter) that drafts 16-32 tokens per pass and verifies them with rollback-free tree verification tailored to Qwen3.5/3.6's Gated DeltaNet layers. On a base M5 MacBook Pro, Uzu generated a 4-bit Qwen3.5 9B response at 92-117 tokens/sec versus 22-25 tokens/sec for llama.cpp and MLX. The piece walks through the CLI and HTTP server setup, explains the metrics the Mirai CLI reports, and provides a standalone Python benchmarking client to compare Uzu, MLX, and llama.cpp under matched conditions.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://blog.dailydoseofds.com/p/karpathys-trick-for-faster-local>

## Questions this post answers

### Why is local LLM decoding at batch size one limited by memory bandwidth instead of compute?

Each decoded token requires streaming every weight of the model from memory into the GPU once, doing roughly one multiply-add per weight, so the GPU finishes computing before the next weight block arrives. This caps throughput at memory bandwidth divided by model size: on a base Apple M5 (153 GB/s) running a 4-bit 9B model (5.2 GB), the hard ceiling is about 29 tokens per second regardless of GPU speed.

_Anyone tuning on-device inference speed can track engineering breakdowns like this one on daily.dev._

### How much faster is the Uzu inference engine than llama.cpp and MLX for local Qwen3.5 9B inference on an Apple M5?

Uzu reached 92 to 117 tokens per second generating a Python function with a 4-bit Qwen3.5 9B model on a base M5 MacBook Pro, compared to only 22 and 25 tokens per second for llama.cpp and MLX respectively in the same side-by-side recording, roughly three to four times faster using the same prompt and bit width.

_Developers picking a local inference engine can compare benchmarks like these on daily.dev before committing._

### How does Uzu's Weaver adapter fix incoherent parallel draft tokens in speculative decoding?

Weaver is a 56.7M-parameter autoregressive adapter that selects tokens sequentially from candidate shortlists proposed in parallel by a DFlash drafter, so each selection can depend on earlier selections instead of being chosen independently. This avoids incoherent outputs like "a few of milk" and instead produces coherent paths like "a gallon of milk," improving acceptance over a plain parallel draft by 24.7 percent in the paper's benchmark.

_Those building speculative decoding pipelines can follow adapter-based fixes like Weaver on daily.dev._

## Similar posts on daily.dev

- [The Infrastructure Behind Making Local LLM Agents Actually Useful](https://daily.dev/posts/the-infrastructure-behind-making-local-llm-agents-actually-useful-ff4qnxir1) · Towards Data Science · 1 upvotes · 0 comments
- [I used speculative decoding to make my local LLM feel instant, and now I actually prefer it to cloud APIs](https://daily.dev/posts/i-used-speculative-decoding-to-make-my-local-llm-feel-instant-and-now-i-actually-prefer-it-to-cloud-eipbr4usd) · XDA Developers · 1 upvotes · 0 comments
- [Local LLMs Are Getting Easier: The Complete Guide \(2026\)](https://daily.dev/posts/local-llms-are-getting-easier-the-complete-guide-2026--gaggzz9j2) · SitePoint · 0 upvotes · 0 comments

---

Tags: [#data-science](https://daily.dev/tags/data-science), [#ai-inference](https://daily.dev/tags/ai-inference), [#llama-cpp](https://daily.dev/tags/llama-cpp)

[View this post on daily.dev](https://daily.dev/posts/karpathy-s-trick-for-faster-local-llms-finally-has-proper-tooling-heblac4j0)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Karpathy’s Trick for Faster Local LLMs Finally Has Proper Tooling","url":"https://daily.dev/posts/karpathy-s-trick-for-faster-local-llms-finally-has-proper-tooling-heblac4j0","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/karpathy-s-trick-for-faster-local-llms-finally-has-proper-tooling-heblac4j0"},"datePublished":"2026-10-08T22:41:41.093Z","dateModified":"2026-10-09T01:08:47.792Z","description":"On-device LLM inference at batch size one is bottlenecked by memory bandwidth rather than compute, since every decoded token requires streaming the full model...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/8e2aef816b7017c5255899de129bbd74?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/8e2aef816b7017c5255899de129bbd74?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Daily Dose of Data Science | Avi Chawla | Substack","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Daily Dose of Data Science | Avi Chawla | Substack","logo":"https://media.daily.dev/image/upload/s--4IHQgTOw--/f_auto/v1710503712/logos/dailydoseofds","url":"https://daily.dev/sources/dailydoseofds"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/karpathy-s-trick-for-faster-local-llms-finally-has-proper-tooling-heblac4j0","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"data-science,ai-inference,llama-cpp","timeRequired":"PT20M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Daily Dose of Data Science | Avi Chawla | Substack","item":"https://daily.dev/sources/dailydoseofds"},{"@type":"ListItem","position":3,"name":"Karpathy’s Trick for Faster Local LLMs Finally Has Proper Tooling"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/karpathy-s-trick-for-faster-local-llms-finally-has-proper-tooling-heblac4j0#faq","mainEntity":[{"@type":"Question","name":"Why is local LLM decoding at batch size one limited by memory bandwidth instead of compute?","acceptedAnswer":{"@type":"Answer","text":"Each decoded token requires streaming every weight of the model from memory into the GPU once, doing roughly one multiply-add per weight, so the GPU finishes computing before the next weight block arrives. This caps throughput at memory bandwidth divided by model size: on a base Apple M5 (153 GB/s) running a 4-bit 9B model (5.2 GB), the hard ceiling is about 29 tokens per second regardless of GPU speed. Anyone tuning on-device inference speed can track engineering breakdowns like this one on daily.dev."}},{"@type":"Question","name":"How much faster is the Uzu inference engine than llama.cpp and MLX for local Qwen3.5 9B inference on an Apple M5?","acceptedAnswer":{"@type":"Answer","text":"Uzu reached 92 to 117 tokens per second generating a Python function with a 4-bit Qwen3.5 9B model on a base M5 MacBook Pro, compared to only 22 and 25 tokens per second for llama.cpp and MLX respectively in the same side-by-side recording, roughly three to four times faster using the same prompt and bit width. Developers picking a local inference engine can compare benchmarks like these on daily.dev before committing."}},{"@type":"Question","name":"How does Uzu's Weaver adapter fix incoherent parallel draft tokens in speculative decoding?","acceptedAnswer":{"@type":"Answer","text":"Weaver is a 56.7M-parameter autoregressive adapter that selects tokens sequentially from candidate shortlists proposed in parallel by a DFlash drafter, so each selection can depend on earlier selections instead of being chosen independently. This avoids incoherent outputs like \"a few of milk\" and instead produces coherent paths like \"a gallon of milk,\" improving acceptance over a plain parallel draft by 24.7 percent in the paper's benchmark. Those building speculative decoding pipelines can follow adapter-based fixes like Weaver on daily.dev."}}]}
```

