Back to News Feed
NVIDIA Blog4d agoIan Finder

Delivering Vera: NVIDIA’s First CPU Built for Agents Is Shipping Now

NVIDIA is officially moving into a new era of infrastructure. Ian Buck, the company’s Vice President of Hyperscale and HPC, has been personally traversing the AI landscape to hand-deliver the first wave of Vera CPU systems. This rollout marks a pivotal moment for the industry, as the hardware begins shipping at scale to the world’s most influential AI labs and cloud providers.

Editor’s Note: This report, originally published on May 18, 2026, has been updated as of Thursday, August 27, 2026, to reflect the latest developments in NVIDIA’s ecosystem.

A Major Expansion in Seattle

The most recent milestone in the Vera rollout involves a significant deepening of the 16-year partnership between NVIDIA and Amazon Web Services (AWS). The two companies announced a massive expansion that includes plans for 2 million additional NVIDIA GPUs and the integration of Vera CPU-based infrastructure into the AWS cloud.

To mark the occasion, Buck hand-delivered the first NVIDIA Vera CPU server and Vera Rubin GPU to the AWS headquarters in Seattle. This delivery follows a rapid, high-profile tour that has already seen Vera systems arrive at Oracle Cloud Infrastructure (OCI) and three of the industry’s most prominent AI research labs: Anthropic, OpenAI, and SpaceXAI.

"Agentic AI is creating a new CPU moment in the AI factory — as models move from answering to acting, Vera is purpose-built to keep that work moving at scale," said Ian Buck.

Why Vera, Why Now?

The shift toward "Agentic AI" has fundamentally changed the requirements for data center hardware. While GPUs handle the heavy lifting of model training and inference, the rise of autonomous agents—which build software, analyze complex datasets, and run simulations—has placed unprecedented pressure on the CPU.

Every orchestration layer, tool call, and long-context retrieval operation is a CPU-intensive task. Traditional processors, often designed for core density rather than real-time, concurrent reasoning, are struggling to keep pace. Vera was engineered specifically to solve this bottleneck.

Key Technical Specifications of the Vera CPU:

  • Core Architecture: 88 custom-designed NVIDIA Olympus cores.
  • Memory Performance: 1.2TB/s of memory bandwidth.
  • Efficiency: Up to 1.8x faster per-core performance specifically for agentic AI workloads.

By completing these complex, real-time tasks more quickly, Vera increases the overall efficiency of the "AI factory," ensuring that users experience faster response times and higher throughput.

Scaling the Cloud: The OCI Deployment

At the Oracle AI Customer Excellence Center, the OCI team—including Product Management lead Karan Batta and Chief Customer and Partner Success Officer Gary Miller—recently unboxed their first Vera systems.

OCI is currently the first cloud provider to deploy Vera at a hyperscale level. According to Batta, the decision to commit to hundreds of thousands of units starting in 2026 was driven by the necessity of sustained performance.

"Vera’s architecture is purpose-built for high-throughput reasoning workloads, delivering the efficiency, density and footprint OCI needs to power the next generation of enterprise AI," Batta noted.

For Oracle’s enterprise clients, this means access to production-grade agentic AI infrastructure that is currently unmatched in the cloud market. Miller emphasized that the center is already preparing to help customers validate their own agent-based workloads using the new hardware.

Hands-On with the AI Labs

The journey of the Vera CPU has been as much about collaboration as it has been about hardware distribution.

  • SpaceXAI: During a visit to the Palo Alto offices in May, the NVIDIA team provided a deep dive into the system’s architecture for Elon Musk. The discussion centered on memory layout, cooling, and the potential for Vera to accelerate reinforcement learning and agent-based simulation pipelines.
  • OpenAI: At the Mission Bay headquarters, Sachin Katti, head of compute infrastructure, received the delivery. In a demonstration of the hardware’s accessibility, Buck used a screwdriver to open the server chassis, showcasing the internal design to the OpenAI team.
  • Anthropic: The first delivery of the spring took place at Anthropic’s SoMa offices. James Bradbury, head of compute, reviewed the Vera motherboard, noting that scaling compute remains a vital accelerant for the growth and capability of their models.

The Future of the AI Factory

Vera is not an isolated product; it is a central component of NVIDIA’s broader "extreme codesign" strategy. It functions alongside the NVIDIA Rubin GPU, the BlueField-4 DPU, the Spectrum-X networking platform, and the MGX rack architecture.

Perhaps most importantly, Vera serves as the host processor for the Vera Rubin NVL72 system. In this configuration, Vera pairs with a pair of Rubin GPUs via second-generation NVIDIA NVLink-C2C. By sharing a unified memory architecture, the system maintains high utilization of accelerated compute, with the CPU handling the orchestration and data movement required to feed the GPUs at double the energy efficiency of traditional server designs.

As the industry transitions from simple chatbots to autonomous agents capable of complex reasoning and execution, the infrastructure must evolve. With Vera now shipping at scale, NVIDIA has provided the engine for this next generation of AI. The age of agentic AI has arrived, and it finally has a CPU built to sustain it.

#agentsnvidia