Hugging Face
Read post

Deploy local agents everywhere with LFM2.5-2.6B

Liquid AI releases LFM2.5-2.6B, a 2.6B parameter model designed for on-device agentic workloads. It supports tool calling and multi-step workflows while running efficiently on consumer hardware — 220 tok/s on Apple M5 Max, 113 tok/s on AMD Ryzen, and even 30 tok/s on phones, using under 2.5 GB of memory. Training involved supervised fine-tuning, multi-domain distillation from specialist teachers, and agentic reinforcement learning inside real agent harnesses. Benchmarks show it competes with models up to 4x its size on instruction following and tool use, with coding being the one area where larger models maintain a clear advantage. The model ships with day-one support for llama.cpp, MLX, vLLM, SGLang, and ONNX, and is available on Hugging Face.

    #llm#ai-agents#local-ai
Aug 04•5m read time•From huggingface.co
Post cover image
Table of contents
How we built a reliable agentic model for edge devicesBenchmark resultsInference speed on CPU and GPUHow to use LFM2.5-2.6BLFM2.5-2.6B demoGet StartedCitation
320 Impressions
Hugging Face's image
Hugging Face

HuggingFace's platform is a resource for developers and researchers working in natural language proc...

639 Followers

•

2.2K Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard