MiniMax M3 is a 428B parameter MoE vision-language model supporting up to 1M token context, native multimodal input (text, image, video), and agentic workflows. It introduces MiniMax Sparse Attention (MSA) for 9x faster prefill and 15x faster decoding compared to M2 at 1M-token context. Deployment options on NVIDIA infrastructure include TensorRT LLM, SGLang, and vLLM with concrete configuration examples. NVIDIA Dynamo enables disaggregated prefill/decode serving for 4x interactivity gains on Blackwell GPUs at 32k input sequence length. Fine-tuning and RL customization are available via NVIDIA NeMo AutoModel and NeMo RL, supporting SFT, LoRA, and reinforcement learning workflows.

4m read timeFrom developer.nvidia.com
Post cover image
Table of contents
Open source inferenceScaling with NVIDIA DynamoCustomize with NVIDIA NeMo FrameworkGet started today
15 Impressions