Articles
News analysisv1

NVIDIA Moves the NVHBM Controller Into the HBM Base Die, Claiming 25% More XPU Area

NVIDIA says the architecture can deliver up to 30% more bandwidth and 15% lower HBM power than standard HBM4E; independent benchmarks, suppliers and production timing remain undisclosed.

John Wood

NVIDIA added NVHBM to NVLink Fusion on August 26. Rather than changing HBM’s role as high-bandwidth memory, the design moves the memory controller from a custom XPU into the HBM stack’s base die, leaving a smaller custom PHY on the XPU. NVIDIA says that, against standard HBM4E, the approach can deliver up to 30% more bandwidth, 15% lower HBM power use and up to 25% more area on the XPU compute die. Those are NVIDIA architectural claims, not independently benchmarked results.

In a conventional design, the controller on the XPU consumes silicon that could otherwise serve compute. Moving it to the base die could leave more XPU area for compute logic. NVIDIA also says it will establish a standard implementation available from multiple memory suppliers, reducing integration and qualification work across vendors. For cloud companies designing their own AI accelerators, that interface and supply-chain proposition may matter more than a standalone bandwidth number.

NVIDIA named Amazon’s Annapurna Labs as the first NVHBM collaborator and connected the effort with Trainium4’s planned NVLink Fusion support. The announcement does not name qualified memory suppliers or disclose sample timing, price, capacity configuration or volume. The collaboration therefore confirms a development direction; it does not establish a new level of HBM procurement, nor does it show that existing GPU racks will use NVHBM.

Tom’s Hardware independently notes that NVHBM is not a replacement for commodity HBM. It is a component for NVLink Fusion custom-chip partners. Extra bandwidth can translate into higher throughput only for memory-bound inference or training work, while system performance still depends on compute, networking, software and cooling. The 15% power figure is likewise NVIDIA’s comparison with standard HBM4E; the public material does not provide a reproducible system-level measurement.

John Wood’s analysis is that NVHBM shifts the competition from stack bandwidth alone toward a joint controller, PHY and multi-supplier qualification design. Its commercial value depends on whether it reduces verification cost for custom XPUs. The next evidence to watch is named memory partners, deployment timing in Trainium4 or other customer chips, and bandwidth and performance results using comparable models, capacity and system-power definitions. Until then, 30% and 15% are vendor targets, not verified market outcomes.

NVIDIA Moves the NVHBM Controller Into the HBM Base Die, Claiming 25% More XPU Area|MemoryTicker