<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/nvidia-s-nemotron-super-why-a-chip-maker-builds-its-own-llms-acut7hwaj" -->

---
title: NVIDIA&#x27;s Nemotron Super: why a chip maker builds its own...
description: NVIDIA released Nemotron 3 Super, a 120B-parameter hybrid model with only 12B active parameters at inference, designed for multi-agent AI workloads. It...
canonical: https://daily.dev/posts/nvidia-s-nemotron-super-why-a-chip-maker-builds-its-own-llms-acut7hwaj
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: NVIDIA&#x27;s Nemotron Super: why a chip maker builds its own LLMs | daily.dev
og:description: NVIDIA released Nemotron 3 Super, a 120B-parameter hybrid model with only 12B active parameters at inference, designed for multi-agent AI workloads. It...
og:url: https://daily.dev/posts/nvidia-s-nemotron-super-why-a-chip-maker-builds-its-own-llms-acut7hwaj
og:image: https://api.daily.dev/og/posts/ACuT7hwAj.png
og:image:alt: NVIDIA&#x27;s Nemotron Super: why a chip maker builds its own LLMs
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# NVIDIA's Nemotron Super: why a chip maker builds its own LLMs

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

NVIDIA released Nemotron 3 Super, a 120B-parameter hybrid model with only 12B active parameters at inference, designed for multi-agent AI workloads. It combines a Mamba-Transformer backbone (linear-time long-sequence handling), Latent MoE (4x more expert consultations at same compute), multi-token prediction for 3x faster inference, a 1M-token context window, and native NVFP4 precision optimized for Blackwell GPUs. Together these deliver up to 5x higher throughput and 2x higher accuracy over the previous Nemotron Super. Pretrained on 25 trillion tokens and fine-tuned with RL across 21 configurations, the model is fully open-sourced with weights, training data, and deployment recipes for vLLM, SGLang, and TensorRT-LLM, available on Hugging Face, NVIDIA NIM, and major cloud providers.

## Content

NVIDIA has released Nemotron 3 Super, a 120-billion-parameter hybrid model with only 12 billion active parameters at inference. It's designed specifically for multi-agent AI workloads, where two problems tend to compound quickly: context explosion (agentic tasks can generate up to 15x more tokens than standard chat) and the "thinking tax" of running full reasoning at every step in a pipeline.

## Architecture

The model combines several techniques to address these problems:

- **Hybrid Mamba-Transformer backbone** - Mamba state space model (SSM) layers handle long sequences in linear time rather than quadratic, giving roughly 4x better memory and compute efficiency compared to pure-transformer approaches.
- **Latent MoE** - A mixture-of-experts design that allows 4x more expert consultations at the same compute cost, improving accuracy without proportionally increasing inference cost.
- **Multi-token prediction (MTP)** - Built-in speculative decoding that delivers up to 3x faster inference without a separate draft model.
- **1-million-token context window** - Handles the long context chains that agentic tasks accumulate.
- **Native NVFP4 precision** - Pretrained at reduced floating-point precision rather than quantized after training, which NVIDIA says preserves more accuracy than post-training quantization. Optimized for Blackwell GPUs.

Combined, these deliver up to 5x higher throughput and 2x higher accuracy over the previous Nemotron Super.

## Training

The model was pretrained on 25 trillion tokens, fine-tuned on 7 million supervised samples, and trained with reinforcement learning across 21 configurations using NeMo Gym environments.

## Why a chip company builds LLMs

NVIDIA VP of Generative AI Kari Briski has been direct about the reasoning: optimizing hardware for AI workloads requires deeply understanding those workloads. The hardware-software co-design loop means NVIDIA needs to run the models themselves to know what to optimize. Nemotron is the practical output of that process.

The roadmap treats models like software libraries - regular release cycles, versioned releases, and eventually community PRs to model architecture.

## Availability

Nemotron 3 Super is fully open: weights, training data, training recipes, gym environments, and deployment cookbooks for vLLM, SGLang, and TensorRT-LLM are all published. It's available on Hugging Face, NVIDIA NIM, and through cloud providers including Google Cloud Vertex AI, AWS Bedrock, Oracle, and Azure.

The open release is notable for enterprise use - companies can fine-tune on their own domain data without the licensing concerns that come with some other frontier models.

---

Tags: [#llm](https://daily.dev/tags/llm), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts), [#nvidia](https://daily.dev/tags/nvidia)

[View this post on daily.dev](https://daily.dev/posts/nvidia-s-nemotron-super-why-a-chip-maker-builds-its-own-llms-acut7hwaj)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"NVIDIA's Nemotron Super: why a chip maker builds its own LLMs","url":"https://daily.dev/posts/nvidia-s-nemotron-super-why-a-chip-maker-builds-its-own-llms-acut7hwaj","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/nvidia-s-nemotron-super-why-a-chip-maker-builds-its-own-llms-acut7hwaj"},"datePublished":"2026-03-11T16:02:36.204Z","dateModified":"2026-04-02T02:16:22.948Z","description":"NVIDIA released Nemotron 3 Super, a 120B-parameter hybrid model with only 12B active parameters at inference, designed for multi-agent AI workloads. It...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/6f8f887d2edbfce91905215f5da1e0c9?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/6f8f887d2edbfce91905215f5da1e0c9?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/nvidia-s-nemotron-super-why-a-chip-maker-builds-its-own-llms-acut7hwaj","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,mixture-of-experts,nvidia","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"NVIDIA's Nemotron Super: why a chip maker builds its own LLMs"}]}
```

