NVIDIA announced that its Groq 3 LPX inference accelerator is now in full production, extending the Vera Rubin NVL72 platform with specialized low-latency token generation for agentic AI workloads. Benchmarked on Gemma 4 31B, it delivered 3,400 output tokens per second for 100,000-token contexts, 4x faster than the nearest alternative. Nebius is first to adopt Groq 3 LPX, CoreWeave has deployed Spectrum-X Multiplane in production, and SpaceXAI is adopting NVIDIA Vera CPUs for agentic AI. NVIDIA also introduced Scale-In, a new accelerated infrastructure class powered by BlueField-4 and DOCA, and detailed NVLink Fusion for connecting custom XPUs to its AI platform. Announcements were made at Hot Chips in Palo Alto.
Table of contents
NVIDIA Partners Adopt Vera Rubin for Lowest Token CostsSpaceXAI Adopts NVIDIA Vera CPUs for Agentic AINVIDIA Groq 3 LPX: The Interactive AI Inference AcceleratorNVIDIA Spectrum-X Multiplane Enables Massive AI Factory Scale on a Flatter, More Resilient NetworkNVIDIA Introduces Scale-In Infrastructure for Agentic AI Factories, Powered by BlueField-4, DOCANVIDIA NVLink Fusion Connects XPUs to NVIDIA’s Leading AI PlatformQuestions this post answers
What token generation speed does NVIDIA Groq 3 LPX achieve on long-context agentic workloads?
NVIDIA Groq 3 LPX delivered 3,400 output tokens per second for 100,000-token long-context use cases in an Artificial Analysis benchmark running the Gemma 4 31B open source agentic model. That figure was reported as 4x faster than the nearest alternative inference platform, positioning it as a low-latency decode accelerator paired with Vera Rubin NVL72 GPUs. Track fast-moving AI inference hardware benchmarks like this one on daily.dev.
How many LP30 accelerators are in a rack-scale NVIDIA Groq 3 LPX deployment?
A rack-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, forming a highly efficient inference engine designed for deterministic, low-latency token generation in modern AI factories running agentic workloads. Follow rack-scale AI infrastructure announcements as they land on daily.dev.
What is NVIDIA Spectrum-X Multiplane and how much bandwidth does it preserve during a plane failure?
NVIDIA Spectrum-X Multiplane splits each server's network connection into several independent paths, or planes, each running its own lightweight two-tier network, allowing Ethernet to scale to 512,000 GPUs without adding a costly third network tier. In an eight-plane topology, if one plane fails, the network still maintains about 90% of total bandwidth, with hardware recovery 11x faster than software-based load balancing. Compare AI factory networking architectures like this one on daily.dev before scaling infrastructure.