With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
The evolution of artificial intelligence is moving beyond the era of singular breakthroughs in chip design or network architecture. Instead, the next frontier is defined by the seamless integration of every component within the modern "AI factory." Recognizing this shift, NVIDIA is significantly expanding the capabilities of its Vera Rubin NVL72 platform, introducing high-speed token generation specifically engineered for the demands of agentic AI systems.
Today, NVIDIA announced that the NVIDIA Groq 3 LPX—a critical component of the Vera Rubin rack-scale system—has officially entered full production. This development marks a pivotal moment for enterprises looking to scale reasoning-heavy workloads. In recent benchmarks conducted by Artificial Analysis using the open-source agentic model Gemma 4 31B, the system achieved an impressive 3,400 output tokens per second. This performance is particularly vital for 100,000-token long-context use cases, outperforming the nearest alternative platform by a factor of four.
The Rise of the Agentic AI Factory
As the industry pivots from model training to active reasoning and agentic workflows, inference has emerged as the new primary challenge. Modern AI agents are not merely generating text; they are processing massive context windows, collaborating with other autonomous systems, and executing complex, multi-step problem-solving tasks.
These workloads necessitate a specialized infrastructure that prioritizes throughput, responsiveness, and economic efficiency at an unprecedented scale. At this week’s Hot Chips conference in Palo Alto, NVIDIA showcased how "extreme codesign"—the practice of architecting compute, networking, and inference acceleration as a unified, cohesive system—is fundamentally reshaping the AI factory.
"Breakthrough performance comes not from optimizing individual components in isolation, but from codesigning every layer of the stack. From networking and context processing to large-scale inference, NVIDIA’s full-stack platform turns AI factories into integrated engines for intelligence."
Strategic Adoption Across the Industry
The Vera Rubin platform is already gaining significant traction among global industry leaders who are looking to future-proof their infrastructure:
- SpaceXAI: The company has announced plans to utilize NVIDIA Vera CPUs to power its next generation of agentic AI, spanning operations from terrestrial data centers to orbital satellite networks.
- CoreWeave: The provider has successfully deployed Spectrum-X Multiplane into production, creating a high-bandwidth, flat, and lossless network architecture that connects Vera Rubin racks via parallel switches.
- Nebius: As the first AI cloud provider to adopt the NVIDIA Groq 3 LPX, Nebius is positioning itself to offer developers industry-leading token generation speeds for highly interactive, real-time AI applications.
NVIDIA Groq 3 LPX: Redefining Inference Latency
Agentic AI introduces a unique performance bottleneck: decode latency. Because agents must reason, utilize external tools, and interact with other systems, they generate responses token by token. Even minor delays in this process can compound, leading to significant latency in complex chains of work.
The NVIDIA Groq 3 LPX is designed to solve this by extending the Vera Rubin NVL72 platform with specialized acceleration for token generation. While Rubin GPUs handle the heavy lifting of large-scale context processing, the LPX architecture accelerates latency-sensitive decode workloads. This synergy ensures that AI factories can deliver responsive reasoning and smoother agent interactions without sacrificing infrastructure efficiency.
Extreme Codesign: The LPU and GPU Synergy
Unlike traditional standalone accelerators, the Groq 3 LPX functions through a deep integration of GPUs and LPUs (Language Processing Units). By computing every layer of an AI model jointly, these components enable a new tier of inference performance. At scale, a rack-scale deployment of Groq 3 LPX can integrate 256 LP30 accelerators connected via direct chip-to-chip links, forming a deterministic, highly efficient inference engine.
Scaling the Network: Spectrum-X Multiplane
As AI factories grow in size, the network often becomes the primary constraint on performance. NVIDIA is addressing this with Spectrum-X Multiplane, an evolution of its Ethernet architecture that allows for massive scaling without the latency, jitter, or prohibitive costs associated with traditional three-tier network designs.
By splitting each server’s network connection into independent "planes," Spectrum-X Multiplane creates a flat, resilient network capable of scaling to 512,000 GPUs.
- Automatic Traffic Management: A hardware engine within the ConnectX SuperNIC manages traffic across these planes, providing instant rerouting in the event of a failure.
- Resilience: In an eight-plane topology, the network retains approximately 90% of its total bandwidth even if one plane fails, with hardware-based recovery speeds 11x faster than software-based alternatives.
- Performance: The architecture delivers 1.6x higher AI networking performance compared to standard off-the-shelf Ethernet solutions.
Introducing Scale-In Infrastructure
To support the operational requirements of these massive factories, NVIDIA is introducing Scale-In, the fifth pillar of its AI networking strategy. Powered by the BlueField-4 processor and the NVIDIA DOCA software platform, Scale-In transforms the traditional access network into a unified, accelerated infrastructure domain.
This new class of infrastructure offloads critical services—such as security, storage access, and real-time observability—from the host compute resources. By accelerating these services in silicon, Scale-In ensures that the infrastructure can scale alongside the AI compute, providing a manageable and secure environment for deploying agentic AI at massive scale.
NVLink Fusion: Bridging Custom Silicon
Recognizing that hyperscalers often require flexibility, NVIDIA has introduced NVLink Fusion. This platform allows partners to integrate custom XPUs and CPUs into the NVIDIA ecosystem.
By utilizing sixth-generation NVLink and NVLink Switch technology, along with NVLink-C2C for energy-efficient connectivity, NVLink Fusion enables companies to build semi-custom AI factories. This approach allows operators to standardize their rack footprints, cooling, and power delivery while maintaining the freedom to innovate with custom silicon.
The Road Ahead
The transition to the agentic era is not merely a software challenge; it is an infrastructure revolution. By integrating the Groq 3 LPX, Spectrum-X Multiplane, and the Scale-In infrastructure, NVIDIA is providing the building blocks for the next generation of AI factories.
As reasoning models continue to grow and agentic workflows generate an ever-increasing volume of tokens, the demand for this purpose-built, high-throughput architecture will only intensify. With the Vera Rubin NVL72 platform at the center, NVIDIA is setting the stage for a future where AI is not just faster, but more responsive, reliable, and economically viable at a global scale.