<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut-iien8k4yn" -->

---
title: NVIDIA Vera Rubin NVL72 Delivers Leading Performance in...
description: NVIDIA&#x27;s Vera Rubin NVL72 system made its debut in MLPerf Inference v6.1 benchmarks, delivering up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and...
canonical: https://daily.dev/posts/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut-iien8k4yn
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut | daily.dev
og:description: NVIDIA&#x27;s Vera Rubin NVL72 system made its debut in MLPerf Inference v6.1 benchmarks, delivering up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and...
og:url: https://daily.dev/posts/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut-iien8k4yn
og:image: https://api.daily.dev/og/posts/IIEN8k4YN.png
og:image:alt: NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

**[NVIDIA](https://daily.dev/sources/nvidia)** · 5 min read · 0 upvotes · 0 comments

## Summary

NVIDIA's Vera Rubin NVL72 system made its debut in MLPerf Inference v6.1 benchmarks, delivering up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1. GB300 NVL72 demonstrated 99% scaling efficiency across a 288-GPU, four-rack configuration, and software optimizations alone drove up to 1.6x performance gains over the prior v6.0 submission. NVIDIA also reported 30x better performance than GB300 NVL72 on SemiAnalysis AgentX agentic benchmarks, and 19 partners including Azure, Dell, Oracle Cloud, and CoreWeave submitted results on the platform.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference>

## Questions this post answers

### How much faster is NVIDIA Vera Rubin NVL72 than GB300 NVL72 in MLPerf Inference benchmarks?

Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on the Qwen3-VL benchmark across offline, server and interactive scenarios using vLLM with NVIDIA Dynamo, and up to 2.5x higher throughput on DeepSeek-R1 using NVIDIA TensorRT-LLM, in its first MLPerf Inference v6.1 preview submission.

_teams comparing GPU generations for inference workloads can follow benchmark shifts like this one on daily.dev._

### What scaling efficiency did NVIDIA GB300 NVL72 achieve when scaling from one rack to four racks in MLPerf Inference v6.1?

NVIDIA's GB300 NVL72 DeepSeek-R1 submission achieved 99% scaling efficiency scaling from a single 72-GPU rack to four racks totaling 288 GPUs in the offline scenario, with throughput growing nearly in proportion to the added hardware, reflecting well-matched architecture, interconnect and software.

_engineers planning multi-rack GPU deployments track scaling efficiency benchmarks like this via daily.dev._

### How much did software optimizations alone improve NVIDIA GB300 NVL72 performance between MLPerf Inference v6.0 and v6.1?

Software optimizations delivered up to 1.6x higher performance on Qwen3-VL for GB300 NVL72 in v6.1 compared with v6.0, driven by lower KV cache precision, additional kernel fusion, improved kernels, and disaggregated serving using vLLM and NVIDIA Dynamo, with further unverified gains reported after the v6.1 submission deadline.

_anyone weighing hardware versus software gains in inference costs can track these updates through daily.dev._

## Similar posts on daily.dev

- [NVIDIA Extreme Co-Design Delivers New MLPerf Inference Records](https://daily.dev/posts/nvidia-extreme-co-design-delivers-new-mlperf-inference-records-jwjjtv0yx) · NVIDIA Developer · 1 upvotes · 0 comments
- [NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt](https://daily.dev/posts/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt-r1iqkqtkn) · NVIDIA Developer · 0 upvotes · 0 comments

---

Tags: [#nvidia](https://daily.dev/tags/nvidia), [#gpu](https://daily.dev/tags/gpu), [#ai-inference](https://daily.dev/tags/ai-inference)

[View this post on daily.dev](https://daily.dev/posts/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut-iien8k4yn)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut","url":"https://daily.dev/posts/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut-iien8k4yn","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut-iien8k4yn"},"datePublished":"2026-09-16T18:38:39.765Z","dateModified":"2026-09-16T20:05:53.481Z","description":"NVIDIA's Vera Rubin NVL72 system made its debut in MLPerf Inference v6.1 benchmarks, delivering up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/f1eb5468062275c5727f8c7923cbcf97?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/f1eb5468062275c5727f8c7923cbcf97?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"NVIDIA","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"NVIDIA","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/236a5c1f410e450b8d696cf3a4bbe374","url":"https://daily.dev/sources/nvidia"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut-iien8k4yn","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"nvidia,gpu,ai-inference","timeRequired":"PT5M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"NVIDIA","item":"https://daily.dev/sources/nvidia"},{"@type":"ListItem","position":3,"name":"NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut-iien8k4yn#faq","mainEntity":[{"@type":"Question","name":"How much faster is NVIDIA Vera Rubin NVL72 than GB300 NVL72 in MLPerf Inference benchmarks?","acceptedAnswer":{"@type":"Answer","text":"Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on the Qwen3-VL benchmark across offline, server and interactive scenarios using vLLM with NVIDIA Dynamo, and up to 2.5x higher throughput on DeepSeek-R1 using NVIDIA TensorRT-LLM, in its first MLPerf Inference v6.1 preview submission. teams comparing GPU generations for inference workloads can follow benchmark shifts like this one on daily.dev."}},{"@type":"Question","name":"What scaling efficiency did NVIDIA GB300 NVL72 achieve when scaling from one rack to four racks in MLPerf Inference v6.1?","acceptedAnswer":{"@type":"Answer","text":"NVIDIA's GB300 NVL72 DeepSeek-R1 submission achieved 99% scaling efficiency scaling from a single 72-GPU rack to four racks totaling 288 GPUs in the offline scenario, with throughput growing nearly in proportion to the added hardware, reflecting well-matched architecture, interconnect and software. engineers planning multi-rack GPU deployments track scaling efficiency benchmarks like this via daily.dev."}},{"@type":"Question","name":"How much did software optimizations alone improve NVIDIA GB300 NVL72 performance between MLPerf Inference v6.0 and v6.1?","acceptedAnswer":{"@type":"Answer","text":"Software optimizations delivered up to 1.6x higher performance on Qwen3-VL for GB300 NVL72 in v6.1 compared with v6.0, driven by lower KV cache precision, additional kernel fusion, improved kernels, and disaggregated serving using vLLM and NVIDIA Dynamo, with further unverified gains reported after the v6.1 submission deadline. anyone weighing hardware versus software gains in inference costs can track these updates through daily.dev."}}]}
```

