<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/kimi-k3-by-moonshot-now-available-on-modal-hxrqajg7q" -->

---
title: Kimi K3 by Moonshot now available on Modal | daily.dev
description: Moonshot&#x27;s Kimi K3, a 2.8 trillion parameter mixture-of-experts multimodal model with a 1M token context window and native vision, is now available on Modal....
canonical: https://daily.dev/posts/kimi-k3-by-moonshot-now-available-on-modal-hxrqajg7q
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Kimi K3 by Moonshot now available on Modal | daily.dev
og:description: Moonshot&#x27;s Kimi K3, a 2.8 trillion parameter mixture-of-experts multimodal model with a 1M token context window and native vision, is now available on Modal....
og:url: https://daily.dev/posts/kimi-k3-by-moonshot-now-available-on-modal-hxrqajg7q
og:image: https://api.daily.dev/og/posts/HxrqAJg7Q.png
og:image:alt: Kimi K3 by Moonshot now available on Modal
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Kimi K3 by Moonshot now available on Modal

**[Modal](https://daily.dev/sources/modal_labs)** · 4 min read · 9 upvotes · 0 comments

## Summary

Moonshot's Kimi K3, a 2.8 trillion parameter mixture-of-experts multimodal model with a 1M token context window and native vision, is now available on Modal. It ranks as the strongest open model on Artificial Analysis's Intelligence Index, fourth overall among 186 models. Modal partnered with Moonshot and vLLM for day-zero support, offering it via a Shared API with token-based pricing and as a dedicated Auto Endpoint. Modal also trained a custom DFlash speculator tuned to K3's architecture, boosting decode throughput from ~50 tokens/sec to over 100 tokens/sec per GPU, with per-user ceilings above 200 tokens/sec. Key architectural innovations include Kimi Delta Attention for efficient long-context handling and Attention Residuals for improved scaling efficiency. The model uses MXFP4 weights and MXFP8 activations from quantization-aware training, enabling broad hardware compatibility.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://modal.com/blog/kimi-k3-by-moonshot-now-available-on-modal>

## Questions this post answers

### What is Kimi K3 by Moonshot and how large is it?

Kimi K3 is a 2.8 trillion parameter mixture-of-experts transformer with 16 of 896 experts active per token, a 1 million token context window, and native vision support. It ranks fourth overall on Artificial Analysis's Intelligence Index, the highest position among open models. It uses Kimi Delta Attention and Attention Residuals for roughly 2.5x the scaling efficiency of its predecessor K2.

_daily.dev helps engineers evaluating frontier open models like Kimi K3 keep pace with new releases._

### How much does the DFlash speculator improve Kimi K3 inference speed on Modal?

A custom-trained DFlash speculator roughly doubles per-user throughput for Kimi K3 on Modal, from about 50 tokens per second without it to about 100 tokens per second with it, and from roughly 0.8 million to 1 million tokens per minute per GPU on agentic workloads, with a per-user ceiling above 200 tokens per second.

_Teams tuning inference speed for large models can track speculative decoding advances via daily.dev._

### Why did Moonshot rewrite prefix caching for Kimi K3's Kimi Delta Attention?

Kimi Delta Attention broke conventional prefix caching, so Moonshot wrote a new implementation and contributed it to vLLM ahead of releasing K3. Moonshot also used quantization-aware training from the SFT stage onward with MXFP4 weights and MXFP8 activations so the 2.8 trillion parameter model could run efficiently on a wide range of hardware.

_daily.dev keeps developers deploying large MoE models current on serving and caching techniques like these._

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-inference](https://daily.dev/tags/ai-inference), [#vllm](https://daily.dev/tags/vllm), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts), [#kimi](https://daily.dev/tags/kimi)

[View this post on daily.dev](https://daily.dev/posts/kimi-k3-by-moonshot-now-available-on-modal-hxrqajg7q)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Kimi K3 by Moonshot now available on Modal","url":"https://daily.dev/posts/kimi-k3-by-moonshot-now-available-on-modal-hxrqajg7q","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/kimi-k3-by-moonshot-now-available-on-modal-hxrqajg7q"},"datePublished":"2026-07-27T15:26:44.804Z","dateModified":"2026-09-14T09:18:04.541Z","description":"Moonshot's Kimi K3, a 2.8 trillion parameter mixture-of-experts multimodal model with a 1M token context window and native vision, is now available on Modal....","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/a8a5a47127b9ddc7923bb8b717ee0a6e?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/a8a5a47127b9ddc7923bb8b717ee0a6e?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Modal","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Modal","logo":"https://media.daily.dev/image/upload/s--HfKZHzC1--/f_auto/v1750946306/logos/modal_labs","url":"https://daily.dev/sources/modal_labs"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/kimi-k3-by-moonshot-now-available-on-modal-hxrqajg7q","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":9},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-inference,vllm,mixture-of-experts,kimi","timeRequired":"PT4M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Modal","item":"https://daily.dev/sources/modal_labs"},{"@type":"ListItem","position":3,"name":"Kimi K3 by Moonshot now available on Modal"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/kimi-k3-by-moonshot-now-available-on-modal-hxrqajg7q#faq","mainEntity":[{"@type":"Question","name":"What is Kimi K3 by Moonshot and how large is it?","acceptedAnswer":{"@type":"Answer","text":"Kimi K3 is a 2.8 trillion parameter mixture-of-experts transformer with 16 of 896 experts active per token, a 1 million token context window, and native vision support. It ranks fourth overall on Artificial Analysis's Intelligence Index, the highest position among open models. It uses Kimi Delta Attention and Attention Residuals for roughly 2.5x the scaling efficiency of its predecessor K2. daily.dev helps engineers evaluating frontier open models like Kimi K3 keep pace with new releases."}},{"@type":"Question","name":"How much does the DFlash speculator improve Kimi K3 inference speed on Modal?","acceptedAnswer":{"@type":"Answer","text":"A custom-trained DFlash speculator roughly doubles per-user throughput for Kimi K3 on Modal, from about 50 tokens per second without it to about 100 tokens per second with it, and from roughly 0.8 million to 1 million tokens per minute per GPU on agentic workloads, with a per-user ceiling above 200 tokens per second. Teams tuning inference speed for large models can track speculative decoding advances via daily.dev."}},{"@type":"Question","name":"Why did Moonshot rewrite prefix caching for Kimi K3's Kimi Delta Attention?","acceptedAnswer":{"@type":"Answer","text":"Kimi Delta Attention broke conventional prefix caching, so Moonshot wrote a new implementation and contributed it to vLLM ahead of releasing K3. Moonshot also used quantization-aware training from the SFT stage onward with MXFP4 weights and MXFP8 activations so the 2.8 trillion parameter model could run efficiently on a wide range of hardware. daily.dev keeps developers deploying large MoE models current on serving and caching techniques like these."}}]}
```

