vLLM is preparing day-0 open-source serving support for Moonshot AI's Kimi K3, a 2.8-trillion-parameter model with a 1-million-token context window, hybrid KDA/full-attention, Attention Residuals, and native vision. Key engineering work includes a redesigned prefix caching system that separates physical block size from prefix-match granularity to handle KDA's recurrent state, fused kernels for KDA prefill/decode and AttnRes, MXFP4 MoE support with SiTU activation on both NVIDIA and AMD hardware, and an optimized MLA module for prefill/decode disaggregation. The post also advocates for a model-announce-then-release-weights workflow to give inference engine teams a stable integration window.
Table of contents
TL;DRKimi K3 at a GlanceA Collaboration Built Over Multiple Kimi GenerationsThe Hardest Part: Prefix Caching for KDAPerformance Work: Removing the New BottlenecksWhat to Expect on Open-Source DayAcknowledgementsOne More Thing: Why the Announcement and Open-Source Release Are Separated15 Impressions