<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/vllm-support-for-nvidia-vera-rubin-nvl72-7-8x-throughput-over-gb200-nvl72-khaqn0xuj" -->

---
title: vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x...
description: vLLM now runs on NVIDIA's next-generation Vera Rubin NVL72 platform with day-0 model support for DeepSeek, Kimi, GLM, and MiniMax, Rubin-tuned FlashInfer...
canonical: https://daily.dev/posts/vllm-support-for-nvidia-vera-rubin-nvl72-7-8x-throughput-over-gb200-nvl72-khaqn0xuj
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72 | daily.dev
og:description: vLLM now runs on NVIDIA's next-generation Vera Rubin NVL72 platform with day-0 model support for DeepSeek, Kimi, GLM, and MiniMax, Rubin-tuned FlashInfer...
og:url: https://daily.dev/posts/vllm-support-for-nvidia-vera-rubin-nvl72-7-8x-throughput-over-gb200-nvl72-khaqn0xuj
og:image: https://api.daily.dev/og/posts/khAQN0Xuj.png
og:image:alt: vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72

**[vLLM](https://daily.dev/sources/vllm)** · 11 min read · 1 upvotes · 0 comments

## Summary

vLLM now runs on NVIDIA's next-generation Vera Rubin NVL72 platform with day-0 model support for DeepSeek, Kimi, GLM, and MiniMax, Rubin-tuned FlashInfer kernels, and locality-aware MoE that exploits CUDA 13.4's locality domains. Early benchmarks show 7.8x throughput per GPU versus GB200 NVL72 on AgentX at matched interactivity, and up to 3.7x higher VLM throughput versus GB300 NVL72 in MLPerf Inference v6.1. Nightly container builds (CUDA 13.4, PyTorch 2.15) are already available, and the team outlines further planned optimizations including full locality domain support, mega kernels, and new kernel integrations for Kimi K3 and DeepSeek-V4.1-Flash.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://vllm.ai/blog/2026-10-09-vera-rubin-preview>

## Questions this post answers

### How much faster is vLLM on NVIDIA Vera Rubin NVL72 compared to GB200 NVL72?

vLLM running MiniMax M3 on Vera Rubin NVL72 delivers up to 7.84x the throughput per GPU versus GB200 NVL72 at matched interactivity on the AgentX benchmark, and 5.18x higher throughput under a 150 TPS latency constraint. Separately, on MLPerf Inference v6.1's VLM benchmark with Qwen3-VL-235B-A22B, Rubin NVL72 delivered up to 3.7x higher throughput than GB300 NVL72.

_Teams benchmarking inference hardware upgrades can follow Vera Rubin performance data as it lands on daily.dev._

### Do existing vLLM Blackwell kernels work on NVIDIA Rubin GPUs without modification?

Yes, because Rubin builds on Blackwell's architecture family with extended tcgen05 tensor core instructions. Although Rubin is a new compile target (sm107), kernels built for the Blackwell target (sm100f) can run on it unmodified, including GEMM-heavy kernels like attention and MoE, giving vLLM day-0 compatibility on Rubin hardware.

_Engineers porting inference stacks to new GPU generations can track kernel compatibility notes like this on daily.dev._

### What is the locality domain feature in CUDA 13.4 and how does it speed up MoE inference?

Locality domains, introduced in CUDA 13.4, let applications place computation and data within the same non-uniform memory access region so SMs read from nearby HBM with higher bandwidth and lower latency. Applied to MoE decode by splitting expert weights column-wise across domains (split-N), this restricts each domain's SMs to local weight shards, yielding about 1.2x average speedups on Rubin for MiniMax M3 MoE layers.

_Developers tuning memory-bound MoE kernels can keep up with locality domain techniques via daily.dev._

## Similar posts on daily.dev

- [NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt](https://daily.dev/posts/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt-r1iqkqtkn) · NVIDIA Developer · 0 upvotes · 0 comments
- [To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72\!](https://daily.dev/posts/to-infinity-and-beyond-thunderkittens-now-on-nvidia-vera-rubin-nvl72--jamovilkr) · Together AI · 0 upvotes · 0 comments
- [vLLM Reaches 25K Total TPS/GPU on Qwen3.5](https://daily.dev/posts/vllm-reaches-25k-total-tps-gpu-on-qwen3-5-hjwogrfb1) · vLLM · 1 upvotes · 0 comments

---

Tags: [#ai-inference](https://daily.dev/tags/ai-inference), [#vllm](https://daily.dev/tags/vllm), [#mixture-of-experts](https://daily.dev/tags/mixture-of-experts)

[View this post on daily.dev](https://daily.dev/posts/vllm-support-for-nvidia-vera-rubin-nvl72-7-8x-throughput-over-gb200-nvl72-khaqn0xuj)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72","url":"https://daily.dev/posts/vllm-support-for-nvidia-vera-rubin-nvl72-7-8x-throughput-over-gb200-nvl72-khaqn0xuj","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/vllm-support-for-nvidia-vera-rubin-nvl72-7-8x-throughput-over-gb200-nvl72-khaqn0xuj"},"datePublished":"2026-10-10T01:49:07.840Z","dateModified":"2026-10-11T02:51:08.983Z","description":"vLLM now runs on NVIDIA's next-generation Vera Rubin NVL72 platform with day-0 model support for DeepSeek, Kimi, GLM, and MiniMax, Rubin-tuned FlashInfer...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/012ba0d0d5c3b56f0ca817a43156dc25?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/012ba0d0d5c3b56f0ca817a43156dc25?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"vLLM","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"vLLM","logo":"https://media.daily.dev/image/upload/s--hTxEuls9--/f_auto/v1744613054/logos/vllm","url":"https://daily.dev/sources/vllm"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/vllm-support-for-nvidia-vera-rubin-nvl72-7-8x-throughput-over-gb200-nvl72-khaqn0xuj","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-inference,vllm,mixture-of-experts","timeRequired":"PT11M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"vLLM","item":"https://daily.dev/sources/vllm"},{"@type":"ListItem","position":3,"name":"vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/vllm-support-for-nvidia-vera-rubin-nvl72-7-8x-throughput-over-gb200-nvl72-khaqn0xuj#faq","mainEntity":[{"@type":"Question","name":"How much faster is vLLM on NVIDIA Vera Rubin NVL72 compared to GB200 NVL72?","acceptedAnswer":{"@type":"Answer","text":"vLLM running MiniMax M3 on Vera Rubin NVL72 delivers up to 7.84x the throughput per GPU versus GB200 NVL72 at matched interactivity on the AgentX benchmark, and 5.18x higher throughput under a 150 TPS latency constraint. Separately, on MLPerf Inference v6.1's VLM benchmark with Qwen3-VL-235B-A22B, Rubin NVL72 delivered up to 3.7x higher throughput than GB300 NVL72. Teams benchmarking inference hardware upgrades can follow Vera Rubin performance data as it lands on daily.dev."}},{"@type":"Question","name":"Do existing vLLM Blackwell kernels work on NVIDIA Rubin GPUs without modification?","acceptedAnswer":{"@type":"Answer","text":"Yes, because Rubin builds on Blackwell's architecture family with extended tcgen05 tensor core instructions. Although Rubin is a new compile target (sm107), kernels built for the Blackwell target (sm100f) can run on it unmodified, including GEMM-heavy kernels like attention and MoE, giving vLLM day-0 compatibility on Rubin hardware. Engineers porting inference stacks to new GPU generations can track kernel compatibility notes like this on daily.dev."}},{"@type":"Question","name":"What is the locality domain feature in CUDA 13.4 and how does it speed up MoE inference?","acceptedAnswer":{"@type":"Answer","text":"Locality domains, introduced in CUDA 13.4, let applications place computation and data within the same non-uniform memory access region so SMs read from nearby HBM with higher bandwidth and lower latency. Applied to MoE decode by splitting expert weights column-wise across domains (split-N), this restricts each domain's SMs to local weight shards, yielding about 1.2x average speedups on Rubin for MiniMax M3 MoE layers. Developers tuning memory-bound MoE kernels can keep up with locality domain techniques via daily.dev."}}]}
```

