NVIDIA's Vera Rubin platform addresses the scale-up challenges of agentic AI inference by combining the Groq 3 LPX accelerator with Vera Rubin NVL72 GPUs. Agentic workloads introduce non-deterministic trajectories, small batch sizes, and extreme low-latency requirements that conventional networking fabrics can't handle economically. The Groq 3 LPX uses three co-designed technologies — high-radix point-to-point C2C links (2.5 TB/s per LPU), compiler-scheduled data movement, and hardware-driven plesiosynchronous timing — to treat thousands of chips as a single deterministic execution surface. Paired with Vera Rubin NVL72 (3,600 PFLOPS, 20.7 TB HBM4) and NVIDIA Dynamo's Attention-FFN Disaggregation, the platform delivers 400 tokens/sec/user on trillion-parameter MoE models with 400K-token context, claiming up to 35x higher throughput per megawatt versus GB200 NVL72.

7m read timeFrom developer.nvidia.com
Post cover image
Table of contents
Why agentic workloads require predictable scale-up networkingHow NVIDIA Groq 3 LPX addresses scale-up challengesHow agentic workloads benefit from LPU C2C
87 Impressions