NVIDIA Nemotron 3 Nano is now available on Together AI. The model uses a hybrid Mamba-Transformer architecture combined with sparse Mixture-of-Experts, activating only ~3B of its 30B parameters per token for efficient inference. It supports a 1M-token context window and is fully open-weights with open training data and recipes. Together AI positions it for production agentic workloads, offering OpenAI-compatible APIs, high throughput, and cost-efficient inference suited for coding assistants, scientific agents, multi-step tool-use pipelines, and long-context enterprise applications.

3m read timeFrom together.ai
Post cover image
Table of contents
NVIDIA Nemotron 3 Nano on Together AIUse casesTry Nemotron 3 Nano