---
title: "With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents"
url: https://daily.dev/posts/with-groq-3-lpx-in-full-production-nvidia-extends-vera-rubin-inference-for-agents-a5hofykfg
source_url: https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion
type: article
source: "NVIDIA"
published: 2026-08-24T18:09:00.928Z
updated: 2026-08-24T18:09:59.793Z
tags: ["nvidia", "gpu", "agentic-ai", "ai-inference"]
reading_time: 10
upvotes: 0
comments: 0
language: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

**[NVIDIA](https://daily.dev/sources/nvidia)** · 10 min read · 0 upvotes · 0 comments

## Summary

NVIDIA announced that its Groq 3 LPX inference accelerator is now in full production, extending the Vera Rubin NVL72 platform with specialized low-latency token generation for agentic AI workloads. Benchmarked on Gemma 4 31B, it delivered 3,400 output tokens per second for 100,000-token contexts, 4x faster than the nearest alternative. Nebius is first to adopt Groq 3 LPX, CoreWeave has deployed Spectrum-X Multiplane in production, and SpaceXAI is adopting NVIDIA Vera CPUs for agentic AI. NVIDIA also introduced Scale-In, a new accelerated infrastructure class powered by BlueField-4 and DOCA, and detailed NVLink Fusion for connecting custom XPUs to its AI platform. Announcements were made at Hot Chips in Palo Alto.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion>

## Questions this post answers

### What token generation speed does NVIDIA Groq 3 LPX achieve on long-context agentic workloads?

NVIDIA Groq 3 LPX delivered 3,400 output tokens per second for 100,000-token long-context use cases in an Artificial Analysis benchmark running the Gemma 4 31B open source agentic model. That figure was reported as 4x faster than the nearest alternative inference platform, positioning it as a low-latency decode accelerator paired with Vera Rubin NVL72 GPUs.

_Track fast-moving AI inference hardware benchmarks like this one on daily.dev._

### How many LP30 accelerators are in a rack-scale NVIDIA Groq 3 LPX deployment?

A rack-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, forming a highly efficient inference engine designed for deterministic, low-latency token generation in modern AI factories running agentic workloads.

_Follow rack-scale AI infrastructure announcements as they land on daily.dev._

### What is NVIDIA Spectrum-X Multiplane and how much bandwidth does it preserve during a plane failure?

NVIDIA Spectrum-X Multiplane splits each server's network connection into several independent paths, or planes, each running its own lightweight two-tier network, allowing Ethernet to scale to 512,000 GPUs without adding a costly third network tier. In an eight-plane topology, if one plane fails, the network still maintains about 90% of total bandwidth, with hardware recovery 11x faster than software-based load balancing.

_Compare AI factory networking architectures like this one on daily.dev before scaling infrastructure._

## Similar posts on daily.dev

- [How the NVIDIA Vera Rubin Platform is Solving Agentic AI’s Scale-Up Problem](https://daily.dev/posts/how-the-nvidia-vera-rubin-platform-is-solving-agentic-ai-s-scale-up-problem-iyn67thas) · NVIDIA Developer · 0 upvotes · 0 comments
- [Inside NVIDIA Groq 3 LPX: The Low-Latency Inference Accelerator for the NVIDIA Vera Rubin Platform](https://daily.dev/posts/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform-6kpvreurx) · NVIDIA Developer · 1 upvotes · 0 comments

---

Tags: [#nvidia](https://daily.dev/tags/nvidia), [#gpu](https://daily.dev/tags/gpu), [#agentic-ai](https://daily.dev/tags/agentic-ai), [#ai-inference](https://daily.dev/tags/ai-inference)

[View this post on daily.dev](https://daily.dev/posts/with-groq-3-lpx-in-full-production-nvidia-extends-vera-rubin-inference-for-agents-a5hofykfg)
