NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
NVIDIA has officially ushered in the era of "gigascale" computing with the launch of its Vera Rubin platform. As production ramps up for the Vera Rubin NVL72, the infrastructure is already being deployed across a prestigious roster of global partners, including CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius. With a supply chain spanning more than 350 factory sites across 30 countries, this represents the most sophisticated and mature rack-scale manufacturing effort in the history of accelerated computing.
The Vera Rubin platform is engineered from the silicon level to the data center grid, specifically optimized to maximize performance per watt while driving down the cost of AI tokens. Early benchmarks from CoreWeave, utilizing the DeepSeek-R1 model, underscore the platform's efficiency: it delivers 10 times the throughput per megawatt compared to the Grace Blackwell NVL72, hitting the critical performance metric for modern, power-constrained AI factories.
Engineering the Future: Extreme Codesign
The secret behind this leap in capability is NVIDIA’s "extreme codesign" philosophy. Rather than relying on off-the-shelf components, the Vera Rubin NVL72 is a unified system comprising seven distinct chips and five specialized rack trays—including the Vera CPU, Groq 3 LPX, Spectrum-6 SPX, and Vera BlueField-4 STX.
At the heart of this architecture sits the NVIDIA Vera CPU. Purpose-built for the agentic era, the Vera CPU features the custom Olympus core, which provides:
- 2x the single-threaded performance of competing chiplet designs.
- 3x the core-to-core bandwidth.
- 40% lower memory latency.
This makes the Vera CPU the most efficient single-threaded processor currently available for the complex, reasoning-heavy workloads that define modern AI agents.
Purpose-Built Networking for AI Factories
To support the massive scale of these systems, NVIDIA has overhauled its networking stack. The sixth-generation NVLink scale-up technology provides more than double the throughput, three times lower latency, and ten times the packet rate of standard Ethernet.
For scale-out operations, the Spectrum-X Ethernet platform integrates 102.4T Spectrum-6 switch systems and 1.6T ConnectX-9 SuperNICs. This combination enables 1.6x higher RDMA bandwidth than traditional Ethernet, utilizing advanced features like adaptive routing, telemetry, and congestion control. Industry leaders such as Tesla, SpaceXAI, and Microsoft are already integrating these switches to accelerate their infrastructure.
Furthermore, NVIDIA’s co-packaged optics—the industry’s first volume-manufactured solution of its kind—offer 5x lower power consumption and 10x higher Mean Time Between Interruption (MTBI) compared to standard pluggable transceivers.
Sustainability and Operational Efficiency
NVIDIA’s commitment to rack-scale design has yielded significant operational improvements. The Vera Rubin NVL72 eliminates the need for internal cables, fans, or hoses within the tray, reducing assembly time from hours to just one minute.
The system is designed for a 45-degree Celsius liquid cooling inlet temperature, allowing for chiller-free, dry-cooler operation. For large-scale AI facilities, this transition to higher-temperature cooling, combined with a closed-loop liquid system, saves millions of gallons of water per megawatt annually.
Empowering Europe’s Open Model Ecosystem
NVIDIA is also positioning Vera Rubin as the bedrock of European AI sovereignty. A newly expanded partnership between Microsoft and Mistral aims to bring frontier AI to the region, combining open European models with secure, customer-controlled cloud environments.
"Europe wants the world’s most capable AI, running under its own laws, close to home, and fully within its control."
Under a multibillion-dollar infrastructure agreement, Mistral is scaling its GPU capacity with thousands of Vera Rubin units. This infrastructure enables governments and regulated industries to deploy AI that adheres to strict data governance and regional autonomy requirements. With Mistral Medium 3.5 and OCR 4 now available in Microsoft Foundry, enterprises can leverage the same tools across public, private, and disconnected cloud environments.
CoreWeave Validates the 10x Performance Leap
CoreWeave has become the first cloud provider to validate the Vera Rubin NVL72 in a live production environment. By running the DeepSeek-R1 benchmark, CoreWeave confirmed a 10x improvement in tokens per second per megawatt over the Grace Blackwell generation.
The Spectrum-X Ethernet SN6600-LD serves as the backbone for these deployments. Featuring a liquid-cooled design, these switches deliver 1.64 Pb/s per rack, effectively doubling the capacity of previous air-cooled generations. This non-blocking, multi-rail fabric ensures that the Vera Rubin NVL72 functions as a single, unified accelerator, removing the bottlenecks typically associated with mixture-of-experts model architectures.
Google Cloud and the Rise of "Superlearners"
Google Cloud has launched its A5X instance powered by Vera Rubin, specifically targeting startups like London-based Ineffable Intelligence. Ineffable is pioneering "superlearner" systems—agents that learn through continuous interaction with simulated environments rather than static datasets.
These reinforcement learning loops require extreme memory bandwidth and low-latency interconnects. The A5X instance, which utilizes NVIDIA ConnectX-9 SuperNICs and Google’s Virgo networking, allows clusters to scale to nearly a million GPUs. This provides the necessary compute for training, tuning, and serving the next generation of physical and agentic AI models.
Orchestrating the Agentic Era
As AI agents take on more complex reasoning and tool-use tasks, the CPU has become a critical bottleneck. Independent benchmarks from DeepInfra demonstrate that the NVIDIA Vera CPU is more than twice as fast at orchestrating these workflows compared to alternative CPUs.
- 1.6x more concurrent AI agents supported at the same quality of service.
- 2.2x faster orchestration speeds.
These gains allow cloud providers to maximize infrastructure utilization, ensuring that as AI agents consume up to 15x more tokens than traditional applications, the underlying hardware remains cost-effective and performant.
Global Availability via Nebius
The Vera Rubin platform is also expanding its reach through Nebius Cloud, which is deploying the hardware across its European and U.S. data centers. By integrating the Vera Rubin NVL72 and Spectrum-6 switches, Nebius is ensuring that AI-native enterprises have immediate access to the latest accelerated computing stack.
As the industry shifts from simple chatbots to complex, reasoning-based agents, the Vera Rubin platform provides the necessary foundation. By unifying seven chips into a single, codesigned system, NVIDIA has created a blueprint for the future of AI—one that prioritizes performance, sustainability, and the lowest possible cost per token.