Moonshot's Kimi K3, a 2.8 trillion parameter mixture-of-experts multimodal model with a 1M token context window and native vision, is now available on Modal. It ranks as the strongest open model on Artificial Analysis's Intelligence Index, fourth overall among 186 models. Modal partnered with Moonshot and vLLM for day-zero support, offering it via a Shared API with token-based pricing and as a dedicated Auto Endpoint. Modal also trained a custom DFlash speculator tuned to K3's architecture, boosting decode throughput from ~50 tokens/sec to over 100 tokens/sec per GPU, with per-user ceilings above 200 tokens/sec. Key architectural innovations include Kimi Delta Attention for efficient long-context handling and Attention Residuals for improved scaling efficiency. The model uses MXFP4 weights and MXFP8 activations from quantization-aware training, enabling broad hardware compatibility.

4m read timeFrom modal.com
Post cover image
Table of contents
Open, frontier, fast. Pick three.Drafting for delta attentionRun Kimi K3 on Modal
25.5K Impressions