<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/minimax-m3-1m-token-multimodal-reasoning-model-now-available-with-nvidia-and-vllm-support-5gcn4m3hd" -->

---
title: MiniMax M3: 1M-token multimodal reasoning model now...
description: MiniMax has released M3, a 428B sparse mixture-of-experts vision-language model supporting up to 1 million tokens of context with text, image, and video...
canonical: https://daily.dev/posts/minimax-m3-1m-token-multimodal-reasoning-model-now-available-with-nvidia-and-vllm-support-5gcn4m3hd
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: MiniMax M3: 1M-token multimodal reasoning model now available with NVIDIA and vLLM support | daily.dev
og:description: MiniMax has released M3, a 428B sparse mixture-of-experts vision-language model supporting up to 1 million tokens of context with text, image, and video...
og:url: https://daily.dev/posts/minimax-m3-1m-token-multimodal-reasoning-model-now-available-with-nvidia-and-vllm-support-5gcn4m3hd
og:image: https://api.daily.dev/og/posts/5Gcn4m3hd.png
og:image:alt: MiniMax M3: 1M-token multimodal reasoning model now available with NVIDIA and vLLM support
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# MiniMax M3: 1M-token multimodal reasoning model now available with NVIDIA and vLLM support

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 2 upvotes · 0 comments

## Summary

MiniMax has released M3, a 428B sparse mixture-of-experts vision-language model supporting up to 1 million tokens of context with text, image, and video inputs. The key innovation is MiniMax Sparse Attention (MSA), which scores 128-token KV blocks and selects only top blocks per query, achieving roughly 9x faster prefill and 15x faster decoding compared to M2. Weights are available on Hugging Face with a free NVIDIA endpoint for testing. Serving is supported via vLLM (day-0 support with specialized kernels), TensorRT-LLM, and SGLang, with benchmarks showing 8,530 tokens/second on B300 hardware. MXFP8 MoE weights, EAGLE3 speculative decoding (~67% acceptance rate), and fine-tuning via NVIDIA NeMo (SFT, LoRA, RL) are also included. Hardware targets include NVIDIA H200, GB200, B300, and AMD MI300/MI350.

## Content

MiniMax has released M3, a 428-billion-parameter mixture-of-experts vision-language model that handles up to one million tokens of context and accepts text, image, and video inputs natively. Weights are on Hugging Face, and NVIDIA is offering a free endpoint for testing.

## What makes M3 different

The headline feature is MiniMax Sparse Attention (MSA). Instead of attending over every token in a million-token window, MSA scores 128-token KV blocks and selects only the top blocks per query. The result, compared to M2 at 1M-token context, is roughly 9x faster prefill and 15x faster decoding. That's the difference between million-token context being a marketing claim and it being usable in practice.

M3 also ships with MXFP8 MoE weights, with DeepGEMM as the backend on Blackwell hardware and Marlin on Hopper. EAGLE3 speculative decoding is supported, hitting around a 67% acceptance rate in testing.

MiniMax also shipped a custom kernel directly on Hugging Face, which is unusual enough that it caught attention on its own.

## Serving options

vLLM added day-0 support for M3. The implementation includes prefill/decode kernels for sparse GQA, KV-block-major prefill scheduling, fused QKNorm+RoPE+KV insert kernels, and CUDA graph handling for speculative decoding. On a B300, benchmarks show 8,530 tokens/second throughput.

Beyond vLLM, M3 can be served via TensorRT-LLM and SGLang. NVIDIA Dynamo supports disaggregated prefill/decode serving, which reportedly delivers around 4x interactivity gains on Blackwell GPUs at 32k input sequence length.

Hardware targets covered so far: NVIDIA H200, GB200, B300, and AMD MI300/MI350.

## Fine-tuning

NVIDIA NeMo AutoModel and NeMo RL support SFT, LoRA, and reinforcement learning workflows for M3 customization.

## What's coming

The vLLM roadmap lists FP8 KV-cache paths, TRTLLM-Gen MoE kernels, context parallelism, and disaggregated serving as upcoming work.

---

Tags: [#llm](https://daily.dev/tags/llm), [#nvidia](https://daily.dev/tags/nvidia), [#multimodal](https://daily.dev/tags/multimodal), [#vllm](https://daily.dev/tags/vllm), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts)

[View this post on daily.dev](https://daily.dev/posts/minimax-m3-1m-token-multimodal-reasoning-model-now-available-with-nvidia-and-vllm-support-5gcn4m3hd)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"MiniMax M3: 1M-token multimodal reasoning model now available with NVIDIA and vLLM support","url":"https://daily.dev/posts/minimax-m3-1m-token-multimodal-reasoning-model-now-available-with-nvidia-and-vllm-support-5gcn4m3hd","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/minimax-m3-1m-token-multimodal-reasoning-model-now-available-with-nvidia-and-vllm-support-5gcn4m3hd"},"datePublished":"2026-06-12T16:18:21.540Z","dateModified":"2026-06-12T17:17:41.105Z","description":"MiniMax has released M3, a 428B sparse mixture-of-experts vision-language model supporting up to 1 million tokens of context with text, image, and video...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/999c8b2c2000655a64ee3174ac99f64b?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/999c8b2c2000655a64ee3174ac99f64b?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/minimax-m3-1m-token-multimodal-reasoning-model-now-available-with-nvidia-and-vllm-support-5gcn4m3hd","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,nvidia,multimodal,vllm,mixture-of-experts","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"MiniMax M3: 1M-token multimodal reasoning model now available with NVIDIA and vLLM support"}]}
```

