NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
As the demand for trillion-parameter AI models and autonomous agents surges, the limitations of traditional hardware architectures are becoming increasingly apparent. Modern AI performance is no longer defined solely by raw compute power; it relies on a cohesive, unified design that integrates memory, networking, storage, and software. To address these evolving requirements, NVIDIA is expanding its NVLink Fusion platform with the introduction of NVHBM, a next-generation high-bandwidth memory technology designed to optimize performance and efficiency for custom XPUs.
Redefining Memory Architecture
In conventional HBM configurations, the memory controller resides on the XPU die, occupying silicon real estate that could otherwise be utilized for compute-heavy tasks. NVIDIA’s NVHBM shifts this paradigm by embedding a custom memory controller directly into the 3D HBM stack. This architectural pivot offers significant technical advantages:
- Increased Bandwidth: Delivers up to 30% higher memory bandwidth.
- Enhanced Efficiency: Reduces HBM power consumption by 15%.
- Optimized Silicon: Frees up to 25% more area on the XPU compute die compared to standard HBM4E.
By standardizing this implementation across multiple memory providers, NVIDIA is streamlining the qualification process for its partners, effectively accelerating the time-to-market for custom AI silicon.
Strategic Collaboration with AWS
A key milestone in this rollout is the deepened partnership with Amazon’s Annapurna Labs. As the first collaborator to adopt NVHBM, Annapurna Labs will integrate this technology alongside the NVLink scale-up architecture to bolster the efficiency of future AWS infrastructure.
"NVHBM represents a new architectural approach to advancing high-bandwidth memory performance and efficiency. We look forward to this technology collaboration to benefit future AWS infrastructure designs." — Nafea Bshara, Vice President of Annapurna Labs at Amazon.
This collaboration extends to the Trainium4 chip, which will support NVLink Fusion, enabling a seamless, rack-scale architecture where Amazon’s custom silicon and NVIDIA GPUs can operate in tandem.
A Vertically Integrated, Open Ecosystem
NVLink Fusion serves as a bridge, allowing partners to connect custom XPUs and CPUs to NVIDIA’s robust rack-scale platform. By leveraging NVIDIA’s ecosystem—including NVLink chiplets, NVLink-C2C, NVLink Switches, and MGX systems—hyperscalers can focus their engineering efforts on core XPU innovation.
Key Takeaways for AI Innovators
- Reduced Risk: Partners utilize a proven, high-performance technology stack for scale-up and scale-out networking.
- Flexibility: The platform supports a broad range of CPU partners, ASIC designers, and system manufacturers.
- Future-Proofing: NVLink Fusion is offered with every generation of NVIDIA’s rack-scale architecture, ensuring that developers have a consistent, reliable path for deploying semi-custom AI infrastructure.
By combining the flexibility of custom chip design with the power of NVIDIA’s established infrastructure, NVLink Fusion and NVHBM are setting a new standard for the next generation of AI-native computing.