NVIDIA NVLink Fusion

Semi-custom AI infrastructure with industry-proven AI scale-up performance and rack-scale architecture.

Overview

Semi-Custom AI Factories with NVLink Fusion

NVIDIA NVLink™ Fusion is the high-bandwidth, low-latency connective technology and IP that enables hyperscalers and AI natives to deploy custom XPUs and CPUs into NVIDIA’s world-leading AI infrastructure platform. Leverage NVIDIA’s proven scale-up and scale-out technology stack and ecosystem, as well as the MGX™ rack-scale architecture, to reduce development complexity, increase performance, and accelerate time to market for semi-custom AI factories. By standardizing on a single unified architecture, NVLink Fusion simplifies operations across the data center, enables flexible reprovisioning of data center capacity, and allows XPU systems to integrate seamlessly with NVIDIA GPU rack-scale systems for heterogeneous compute. 

NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory

Amazon’s Annapurna Labs will be the first to work with NVIDIA on NVHBM, combining AWS custom silicon with NVIDIA memory technology and the NVLink scale-up architecture to enhance performance and efficiency for AI workloads.

Integrating Semi-Custom Compute into Rack-Scale Architecture with NVIDIA NVLink Fusion

Learn how NVIDIA NVLink Fusion allows hyperscalers to build semi-custom AI infrastructure, integrating their ASICs or CPUs with NVIDIA GPUs, while standardizing on a single scalable hardware infrastructure.

Using NVLink Fusion, high-performance AI factories can scale quickly, benefiting from all solution components that make the NVIDIA rack-scale architecture.

Benefits

NVLink Fusion Benefits

World-Class Scale-Up Performance

Unlocking the full potential of AI factories requires fast memory and swift, seamless communication among all accelerators. NVIDIA NVHBM increases memory bandwidth and reduces power, while NVIDIA NVLink 6 connects 72 XPUs all-to-all at 3.6 TB/s per XPU, with future roadmap domain sizes up to 1,152, to boost AI performance and return on investment.

Lower Development Costs

The established NVLink Fusion supplier ecosystem provides all of the components required for full rack-scale deployment based on the OCP MGX architecture, from rack and chassis to power delivery and cooling systems, eliminating the development costs and deployment risks associated with a new rack design.

Accelerated Time to Market

By leveraging NVIDIA’s battle-tested technology stack and ecosystem of ASIC designers, CPU and IP providers, and OEM/ODMs, hyperscalers can achieve faster time to market and faster time to revenue.

Single Unified Architecture

As hyperscalers already deploy full NVIDIA rack solutions, NVLink Fusion allows heterogeneous silicon offerings while standardizing around a common rack design, accelerating AI factory deployment, and simplifying management.

Platform

NVIDIA NVLink Fusion Platform

NVIDIA NVLink

NVIDIA NVLink 6 and NVLink Switch Chip enable 260 TB/s of bandwidth in a single 72-accelerator NVLink domain (NVL72) and deliver 4x bandwidth efficiency with NVIDIA Scalable Hierarchical Aggregation and Reduction Protocol (SHARP)™ FP8 support.

NVIDIA NVLink-C2C

NVIDIA NVLink-C2C extends the industry-leading NVLink technology to a chip-to-chip interconnect. This enables the creation of a new class of integrated products with NVIDIA partners, built via chiplets, allowing NVIDIA GPUs or CPUs to have a high-bandwidth coherent connection with custom silicon.

NVIDIA NVHBM

NVIDIA NVHBM is a custom next-generation high-bandwidth memory (HBM) architecture. It includes a custom HBM base die and memory controller designed by NVIDIA and validated with leading memory providers to increase bandwidth, reduce HBM power, and save area on the XPU compute die.

AI Infrastructure Platform

NVIDIA provides a modular portfolio of AI factory technology, including NVIDIA GPUs, NVIDIA Vera CPUs, co-packaged optics (CPO) switches, ConnectX® SuperNICs™, BlueField® DPUs, and Mission Control™ software for optimizing AI workflows and managing AI infrastructure.

Full-rack solutions are also available for semi-custom AI factory integration, including the Vera Rubin NVL72 rack, which can be mixed with XPU-based systems for disaggregated inference, the Vera CPU rack for supporting agentic AI systems and reinforcement learning, the NVIDIA LPX rack for assisting high-context and low-latency inference, the NVIDIA STX rack for AI-native storage, and the NVIDIA SPX rack for scale-out networking.

Adopters

NVLink Fusion Ecosystem

Scaling AI Inference Performance with NVLink Fusion

Learn how NVIDIA NVLink Fusion addresses the growing demands of complex AI models.