<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/llm-inference-server-with-continuous-batching-ssd-caching-for-apple-silicon-managed-from-the-mac-wa8k7kkkt" -->

---
title: LLM inference server with continuous batching &amp; SSD...
description: oMLX is an open-source LLM inference server built specifically for Apple Silicon Macs, featuring continuous batching and a tiered KV cache system that spans...
canonical: https://daily.dev/posts/llm-inference-server-with-continuous-batching-ssd-caching-for-apple-silicon-managed-from-the-mac-wa8k7kkkt
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: LLM inference server with continuous batching &amp; SSD caching for Apple Silicon — managed from the macOS menu bar | daily.dev
og:description: oMLX is an open-source LLM inference server built specifically for Apple Silicon Macs, featuring continuous batching and a tiered KV cache system that spans...
og:url: https://daily.dev/posts/llm-inference-server-with-continuous-batching-ssd-caching-for-apple-silicon-managed-from-the-mac-wa8k7kkkt
og:image: https://api.daily.dev/og/posts/WA8k7kkkt.png
og:image:alt: LLM inference server with continuous batching &amp; SSD caching for Apple Silicon — managed from the macOS menu bar
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

**[Jay](https://daily.dev/sources/qwnhvce4dlvpzk8hnrowl)** · [@finallyjay](https://daily.dev/finallyjay) · 7 upvotes · 0 comments

## Summary

oMLX is an open-source LLM inference server built specifically for Apple Silicon Macs, featuring continuous batching and a tiered KV cache system that spans hot RAM and cold SSD storage. It supports text LLMs, vision-language models, OCR, embeddings, and rerankers, all managed through a native macOS menu bar app or web admin dashboard. Key features include multi-model serving with LRU eviction and model pinning, OpenAI/Anthropic API compatibility, MCP tool integration, Claude Code optimization, and one-click integrations with tools like OpenCode and Codex. It can be installed via a .dmg, Homebrew, or from source, and requires macOS 15+ with Apple Silicon (M1–M4).

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://github.com/jundot/omlx>

## Similar posts on daily.dev

- [Don’t just attend KubeCon \+ CloudNativeCon, Merge Forward your experience\!](https://daily.dev/posts/don-t-just-attend-kubecon-cloudnativecon-merge-forward-your-experience--l0rpp73x8) · CNCF · 1 upvotes · 0 comments
- [Announcing H2 2026 KCDs](https://daily.dev/posts/announcing-h2-2026-kcds-m96goajm1) · CNCF · 1 upvotes · 0 comments
- [Two months of Open Community Groups](https://daily.dev/posts/two-months-of-open-community-groups-asf52zhbs) · CNCF · 0 upvotes · 0 comments
- [CNCF Unveils Schedule for KubeCon \+ CloudNativeCon Europe 2026](https://daily.dev/posts/cncf-unveils-schedule-for-kubecon-cloudnativecon-europe-2026-ikhcoa5cb) · CNCF · 2 upvotes · 0 comments
- [CNCF Debuts KubeCon \+ CloudNativeCon Japan 2026 Schedule](https://daily.dev/posts/cncf-debuts-kubecon-cloudnativecon-japan-2026-schedule-xp5pyudub) · CNCF · 1 upvotes · 0 comments

---

Tags: [#python](https://daily.dev/tags/python), [#llm](https://daily.dev/tags/llm), [#mac](https://daily.dev/tags/mac)

[View this post on daily.dev](https://daily.dev/posts/llm-inference-server-with-continuous-batching-ssd-caching-for-apple-silicon-managed-from-the-mac-wa8k7kkkt)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"DiscussionForumPosting","mainEntityOfPage":"https://daily.dev/posts/llm-inference-server-with-continuous-batching-ssd-caching-for-apple-silicon-managed-from-the-mac-wa8k7kkkt","headline":"LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar","text":"Shared: GitHub - jundot/omlx: LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar","url":"https://daily.dev/posts/llm-inference-server-with-continuous-batching-ssd-caching-for-apple-silicon-managed-from-the-mac-wa8k7kkkt","datePublished":"2026-04-29T13:05:04.582Z","dateModified":"2026-07-21T12:18:34.119Z","author":{"@type":"Person","name":"Jay","url":"https://daily.dev/finallyjay","image":"https://lh3.googleusercontent.com/a/ACg8ocL5i3hkxSSWLLoubyZSkPBN6T_N7QRlwpPOOyQWeyJV53Q=s96-c","description":"👨‍💻 Software Engineer","interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"EndorseAction"},"userInteractionCount":2680}},"image":"https://media.daily.dev/image/upload/s--CxzD6vbw--/f_auto/v1722860399/public/Placeholder%2005","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":7},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"sharedContent":{"@type":"WebPage","url":"https://api.daily.dev/r/nFFblB2pF"},"isPartOf":{"@type":"WebPage","url":"https://daily.dev/sources/qwnhvce4dlvpzk8hnrowl","name":"Jay"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Jay","item":"https://daily.dev/sources/qwnhvce4dlvpzk8hnrowl"},{"@type":"ListItem","position":3,"name":"LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar"}]}
```

