vLLM
Read post

A Preview of Production-Scale Kimi K3 Support on vLLM

vLLM is preparing day-0 open-source serving support for Moonshot AI's Kimi K3, a 2.8-trillion-parameter model with a 1-million-token context window, hybrid KDA/full-attention, Attention Residuals, and native vision. Key engineering work includes a redesigned prefix caching system that separates physical block size from prefix-match granularity to handle KDA's recurrent state, fused kernels for KDA prefill/decode and AttnRes, MXFP4 MoE support with SiTU activation on both NVIDIA and AMD hardware, and an optimized MLA module for prefill/decode disaggregation. The post also advocates for a model-announce-then-release-weights workflow to give inference engine teams a stable integration window.

    #ai-inference#vllm#mixture-of-experts#kimi-k3
Jul 22•13m read time•From vllm.ai
Post cover image
Table of contents
TL;DRKimi K3 at a GlanceA Collaboration Built Over Multiple Kimi GenerationsThe Hardest Part: Prefix Caching for KDAPerformance Work: Removing the New BottlenecksWhat to Expect on Open-Source DayAcknowledgementsOne More Thing: Why the Announcement and Open-Source Release Are Separated
15 Impressions
vLLM's image
vLLM

76 Followers

•

163 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard