finallyjay's profile
Jay

@finallyjayβ€’Apr 29
2.7K
Post cover image

GitHub - jundot/omlx: LLM inference server with continuous batching & SSD caching for Apple Silicon β€” managed from the macOS menu bar

From github.comβ€’Apr 24β€’10m read time

oMLX is an open-source LLM inference server built specifically for Apple Silicon Macs, featuring continuous batching and a tiered KV cache system that spans hot RAM and cold SSD storage. It supports text LLMs, vision-language models, OCR, embeddings, and rerankers, all managed through a native macOS menu bar app or web admin dashboard. Key features include multi-model serving with LRU eviction and model pinning, OpenAI/Anthropic API compatibility, MCP tool integration, Claude Code optimization, and one-click integrations with tools like OpenCode and Codex. It can be installed via a .dmg, Homebrew, or from source, and requires macOS 15+ with Apple Silicon (M1–M4).

5.6K Impressions
finallyjay's user avatar
Jay
@finallyjay
JoinedΒ Nov 26. 2023
2.7K

πŸ‘¨β€πŸ’» Software Engineer

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • Β© 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard