NVIDIA has released Nemotron 3 Ultra, a 550B-parameter Mixture-of-Experts model with 55B active parameters designed for long-running agentic workflows. Key architectural innovations include a Hybrid Mamba-Transformer design for long-context efficiency, NVFP4 quantization delivering 5x throughput over BF16 on Blackwell GPUs, LatentMoE for efficient expert routing, and multi-token prediction. A novel Multi-Teacher On-Policy Distillation (MOPD) training method uses 10+ specialized teacher models to improve cross-domain reasoning. The model achieves SWEBench Verified scores of 65–70.4% and reduces agentic task costs by up to 30%. The release also includes Nemotron 3.5 Content Safety (4B guardrail model covering 23 safety categories) and Nemotron 3.5 ASR (multilingual streaming with sub-100ms latency). All weights, data, and training recipes are open under the OpenMDW-1.1 license, with deployment available across AWS, Google Cloud, Azure, and many other platforms.