NVIDIA GB200 Superchip Guide: Liquid-Cooled Racks and Servers

Updated September 2026. This guide now covers the NVIDIA GB200 superchip, its NVL72 and NVL36x2 liquid-cooled racks, and how the platform fits the current GB300 Blackwell Ultra generation.

The NVIDIA GB200 superchip is a key building block of the Blackwell platform. It combines one NVIDIA Grace CPU with two NVIDIA Blackwell B200 Tensor Core GPUs through high-speed NVLink-C2C interconnect technology, creating a tightly integrated accelerated computing module designed for large-scale AI training and inference workloads.

GB200 represents a major shift in AI data center architecture. Instead of scaling performance only by adding more individual servers, NVIDIA introduced rack-scale systems such as the GB200 NVL72, which integrates 36 Grace CPUs and 72 Blackwell GPUs into a single liquid-cooled rack with a large NVLink domain.

This guide explains what the NVIDIA GB200 superchip is, how the GB200 NVL72 and NVL36x2 systems are structured, why liquid cooling is required, and how networking components such as InfiniBand, Ethernet, optical transceivers, and high-speed cables connect these AI clusters.

You will also learn how GB200 compares with the newer GB300 Blackwell Ultra platform and how optical connectivity requirements are evolving toward 800G and beyond.

What Is the NVIDIA GB200 Superchip

What Is the NVIDIA GB200 Superchip?

The GB200 Grace Blackwell superchip brings one Grace CPU and two Blackwell B200 GPUs together on a single board. NVIDIA interconnects the CPU and GPUs over NVLink-C2C at 900 GB/s per link, so the three chips operate as one tightly coupled unit instead of three separate devices.

ComponentGB200 Grace Blackwell Superchip
Configuration1 Grace CPU + 2 Blackwell B200 GPUs
CPU72-core Arm Neoverse V2 (Grace), up to 480 GB LPDDR5X
GPU architectureBlackwell (custom TSMC 4NP process, ~208 billion transistors)
CPU-GPU interconnectNVLink-C2C, 900 GB/s per link
GPU memoryHigh-bandwidth HBM3e per GPU
Board power~2,700 W per superchip
Fabric scaleUp to 72 GPUs in one NVLink domain via fifth-generation NVLink

The GB200 superchip itself is only the compute building block. The real deployment platform is the rack-scale GB200 system, especially the GB200 NVL72, which connects 72 Blackwell GPUs and 36 Grace CPUs into one large NVLink domain.

NVIDIA highlights significant performance improvements compared with previous-generation Hopper systems, including major gains for trillion-parameter AI model inference and training workloads. These figures depend on workload configuration and should be considered platform-level performance comparisons rather than universal benchmarks.

GH200

GH200 vs GB200: What Changed

Comparing the GB200 to its direct predecessor, the GH200, makes the shift clear. The GH200 paired one Grace CPU with one Hopper H200 GPU in a 1:1 ratio. The GB200 pairs one Grace CPU with two Blackwell GPUs in a 1:2 ratio, doubling GPU density per CPU.

FeatureGH200 Grace HopperGB200 Grace Blackwell
CPU to GPU ratio1 Grace: 1 H2001 Grace: 2 B200
GPU memoryUp to 144 GB HBM3e per H200Larger HBM3e pool per B200
CPU-GPU interconnectNVLink-C2C 900 GB/sNVLink-C2C 900 GB/s
Superchip power~1,000 W~2,700 W
CoolingAir-coolable at lower densityLiquid cooling above ~30 kW/rack

The higher power of the GB200 superchip is the reason this platform looks different in the data center. A GB200 NVL72 rack draws roughly 120 kW, while a typical air-cooled H100 rack runs around 40 kW. Once a single rack exceeds roughly 30 kW, air cooling stops being practical, and direct liquid cooling becomes the design baseline rather than an option.

GB200 Rack Systems: NVL72 and NVL36x2

NVIDIA does not deploy the GB200 superchip as an individual server component. Instead, it is delivered through rack-scale systems designed for large AI clusters.

The two main GB200 rack configurations are:

  • GB200 NVL72 — a single liquid-cooled rack containing 72 Blackwell GPUs in one NVLink domain
  • GB200 NVL36x2 — two connected racks that together provide a 72-GPU NVLink domain

Both architectures are designed to provide massive GPU-to-GPU bandwidth for large language models, generative AI, and high-performance computing workloads.

GB200 NVL72

The GB200 NVL72 is NVIDIA’s flagship rack-scale AI system. It integrates:

  • 36 Grace CPUs
  • 72 Blackwell B200 GPUs
  • 9 NVLink Switch trays
  • Liquid cooling infrastructure
  • High-speed scale-out networking

The entire rack forms a single 72-GPU NVLink domain, allowing GPUs to communicate as a unified accelerated computing system rather than as independent servers.

GB200 Rack Systems

The internal architecture includes:

  • 18 × 1RU compute trays 
  • 9 × 1RU NVLink switch trays 
  • Power shelves and rack-level liquid cooling components

Each compute tray contains two GB200 superchips. Since each GB200 superchip includes one Grace CPU and two B200 GPUs, each compute tray provides:

  • 2 Grace CPUs
  • 4 Blackwell GPUs

Across 18 compute trays, the rack delivers:

  • 36 Grace CPUs
  • 72 Blackwell GPUs

The NVLink switch trays provide the high-bandwidth GPU fabric. Each NVLink switch tray contains two NVSwitch ASICs, enabling all GPUs inside the rack to participate in the same NVLink domain.

This architecture allows the GB200 NVL72 rack to function as a large-scale AI computing unit optimized for trillion-parameter model training and inference.

GB200 NVL36x2

The NVL36x2 spreads the same 72-GPU domain across two side-by-side racks. Each rack carries 18 Grace CPUs and 36 Blackwell GPUs in 2U compute nodes, and the two racks interconnect through a non-blocking fabric so all 72 GPUs still communicate as one domain.

Because the distance between the two racks is longer than an in-rack copper run, the NVSwitch trays in an NVL36x2 use 1.6T dual-port OSFP cages to connect the pair. This is the main place where rack-to-rack optical connectivity enters the GB200 picture.

Deployment Reality in 2026

Early coverage, including this article’s 2024 version, described the NVL72 as rarely deployed because most data centers could not support 120 kW per rack. That changed quickly. By 2026 the NVL72 was shipping in production at CoreWeave, Microsoft Azure, Oracle Cloud, and Google Cloud, and Meta built its Catalina pod around the NVL72 architecture.

Two constraints still shape real deployments:

  • Liquid-cooling retrofit backlog. Fewer than 40% of data centers have finished liquid-cooling retrofits, and retrofits take 6 to 12 months. Many orders therefore ship before the facilities are ready.
  • Power per rack. Facilities already running 120 kW racks need headroom for the ~132 kW that GB300 racks draw. This is why the NVL36x2 form factor remains common: it lets operators use two lower-power racks instead of one very high-power rack.

Meta’s Catalina approach shows how operators adapt. Meta runs 120 kW liquid-cooled compute racks inside legacy 20 kW air-cooled data centers by adding air-assisted liquid-cooling (ALC) side pods, as Data Center Dynamics reported.

Liquid Cooling Requirements for GB200

Liquid cooling is not optional for GB200 racks; it is a design requirement. The compute nodes dissipate 5.4 to 5.7 kW each, and most of that heat leaves through direct-to-chip (DTC) cold plates rather than server fans.

Liquid Cooling Requirements for GB200

For data center operators, the practical questions are usually about the facility, not the rack:

  • Coolant loop. GB200 racks connect to a coolant distribution unit (CDU) that supplies the rack loop. NVIDIA cites coolant flow around 2 liters per second into the cabinet.
  • Rack power planning. Plan for ~120 kW per NVL72 rack, ~66 kW per NVL36x2 rack, and ~132 kW per GB300 NVL72 rack.
  • Air-cooled legacy sites. If your facility is air-cooled, an ALC approach like Meta’s Catalina pod can add liquid cooling without a full data center rebuild.

Choosing compatible optical modules and cables for the liquid-cooled rack is a separate decision, covered in the next section.

Networking a GB200 Rack: Cables, Transceivers, and NICs

The GB200 platform is dense with optics, and that is where FiberMall’s 800G OSFP transceiver line fits. The networking picture splits into three tiers.

Compute Fabric (InfiniBand or Ethernet)

Each NVL72 compute node carries four InfiniBand network cards. First-generation GB200 systems ship with NVIDIA ConnectX-7 400G InfiniBand cards, and GB300 moves to ConnectX-8 at 800G. On the InfiniBand side, NVIDIA builds these clusters around the Quantum-X800 platform; on the Ethernet side, the Spectrum-X800 platform plays the same role. NVIDIA explains the full system layout in its multi-node NVLink system guide.

If you are building or upgrading a GB200 cluster, you will be selecting:

  • NVIDIA-compatible ConnectX-7 400G InfiniBand network cards
  • 400G QSFP-DD transceivers for the front-panel ports
  • Quantum-X800 or Spectrum-X800 top-of-rack switches

Rack-to-Rack NVLink (NVL36x2)

The two racks in an NVL36x2 pair connect through 1.6T dual-port OSFP cages on the NVSwitch trays. Because these links carry NVLink traffic, they need high-performance OSFP optics and low-loss cabling. NVIDIA-compatible 800G OSFP modules, such as the SR8 and DR8 families, are the practical building blocks here, and they pair with 800G twin-port OSFP DAC cables for shorter rack-to-rack runs.

One useful rule for planning: keep copper where distances are short and power budgets are tight, and switch to optics where reach or density demands it. Inside a single NVL72 rack, copper NVLink is the right call, which is why NVIDIA saves about 20 kW per cabinet that way. Between the two NVL36x2 racks, and for any rack-to-rack or scale-out link beyond a few meters, optics win. This split is exactly why the GB200 ecosystem uses both DAC cables and OSFP optical modules.

Rack-to-Rack NVLink (NVL36x2)

Management, Storage, and Scale-Out

Beyond the compute fabric, GB200 racks carry management and storage traffic handled by BlueField-3 DPUs and the on-board switch management plane. These links are lower speed and can use 25G, 100G, or 200G modules depending on the switch. NVIDIA also runs storage and management over standard Ethernet within the rack.

Practical Optics Selection for GB200

Link typeTypical portRecommended module
Compute fabric (ConnectX-7)400G QSFP-DD400G SR4 / DR4 transceivers
Compute fabric (ConnectX-8, GB300)800G OSFP800G SR8 / DR8 OSFP modules
NVL36x2 rack-to-rack NVLink1.6T OSFP cage800G twin-port OSFP DAC or optics
Management / storage25G / 100G / 200GSFP28 / QSFP28 / QSFP56 modules
Networking a GB200 Rack

Most GB200 racks run NVIDIA-branded or NVIDIA-compatible modules, so compatibility matters. A compatible 800G OSFP module must match the NVIDIA port’s speed, reach, and optical interface, and it should carry proper digital diagnostics so the switch can monitor link health. FiberMall tests its NVIDIA-compatible modules against these requirements, which is why the 800G OSFP SR8 and DR8 families used across GB200 and GB300 links are built to MSA standards.

If you are planning a GB200 deployment, our guide to NVIDIA’s Spectrum-X solution explains how the Ethernet side of these clusters is architected.

Liquid-Cooled GB200 Servers and Cabinets by Manufacturer

Beyond NVIDIA’s own DGX GB200 NVL72, several server makers build GB200 liquid-cooled racks. The table below summarizes the systems each vendor showed and the design choices that matter.

ManufacturerGB200 systemsKey design notes
NVIDIA (DGX)DGX GB200 NVL72Reference 120 kW rack, 18 x 1U nodes, 9 NVSwitch trays, copper NVLink backplane
SupermicroGB200 NVL72 and NVL36x2 systemsFull liquid-cooled rack options across the MGX portfolio, 1U and 2U nodes
Foxconn / IngrasysDGX GB200 NVL72, NVL36x2Mass production of 72-GPU and 36-GPU rack systems; large order volumes
QCT (Quanta Cloud Technology)QuantaGrid D75B-1U1U node with two GB200 superchips, 8x E1.S SSDs, cold-plate liquid cooling
WiwynnGB200 NVL72 rack systemsRack-level liquid-cooled AI server racks plus the UMS100 cooling management system
ASUSESC AI POD (GB200 NVL72)1U node with dual liquid-cooled GB200 nodes and a 48V-to-12V PDB
InventecArtemis 1U and 2U serversCabinet-level GB200 NVL72 with bus-bar power and side-car cooling cabinets

Detailed Vendor Designs

These details come from each vendor’s GB200 unveilings and help you compare designs before you buy.

Foxconn (Ingrasys). At GTC 2024, Foxconn subsidiary Ingrasys showed the NVL72 liquid-cooled server, which integrates 72 Blackwell GPUs and 36 Grace CPUs interconnected over fifth-generation NVLink. Foxconn moved its DGX GB200 systems into rack-form mass production in the second half of 2024, offering NVL72, NVL36x2, and HGX B200 cabinets, and it became one of the largest winners of the GB200 platform ramp.

Supermicro. Supermicro was an early supporter of the platform, offering GB200 NVL72 and NVL36x2 rack systems alongside its established NVIDIA MGX and GH200 Grace Hopper server lines. Its pitch centers on full liquid-cooled racks and a 100% liquid-cooled data center reference, which suits operators building dense AI clusters.

QCT (Quanta Cloud Technology). QCT’s QuantaGrid D75B-1U is a 1U node built for the GB200 NVL72 framework, fitting up to 72 devices in a single cabinet. Each node carries two GB200 superchips with cold-plate liquid cooling, 480 GB of LPDDR5X CPU memory, eight E1.S SSDs, one M. 2 drive, and PCIe 5.0 expansion slots.

Wiwynn. Wiwynn was one of the first vendors to align with the GB200 NVL72 rack standard. It offers rack-level liquid-cooled AI server racks and the UMS100 universal liquid-cooling management system, which provides real-time monitoring, cooling energy optimization, rapid leak detection, and integration with existing management systems over the Redfish interface.

ASUS. ASUS’s ESC AI POD targets the GB200 NVL72 configuration. Its 1U compute node pairs a bus power supply with two liquid-cooled GB200 nodes, uses a power distribution board to step 48 V down to 12 V for the Blackwell GPUs, and includes E1.S storage plus BlueField-3 DPUs. ASUS also offers the ESC NM2-E1, a dual-GH200 platform for lower-cost Arm-based AI work.

Inventec. Inventec’s Artemis line covers 1U and 2U GB200 servers, and its cabinet-level NVL72 design runs at 120 kW with a 1,400 A bus bar, eight 33 kW power shelves in a 1+1 redundant arrangement, and blind-mate liquid, power, and communication connectors. A rear “Side Car” cabinet handles the liquid-cooling loop.

A few patterns repeat across every vendor:

  • 1U and 2U compute nodes with two Bianca boards each, using cold plates on the GPUs.
  • Shared liquid-cooling loops across compute and NVSwitch trays.
  • Bus-bar power distribution with blind-mate connectors, which is why serviceable racks need careful alignment.
  • Storage on E1.S drives inside each node, kept compact to preserve airflow.

If you are comparing vendors, focus less on the GPU count and more on the networking and cooling specifics: how many front-panel ports each node exposes, whether the NVSwitch interconnects use copper or optics, and how the vendor handles the coolant distribution within the cabinet.

The 2026 Roadmap: GB200 to GB300 and Vera Rubin

NVIDIA has already moved past GB200. The GB300, based on Blackwell Ultra GPUs, is the volume platform for 2026. Industry analysts project up to roughly 60,000 Blackwell Ultra racks shipping in 2026, driven by Microsoft, Amazon, and Meta.

 GB200 NVL72GB300 NVL72 (Blackwell Ultra)
GPU72x B20072x B300
GPU power1,000 W per GPU1,400 W per GPU
Rack power~120 kW~132 kW
GPU memoryHBM3e per B200HBM3e per B300 (288 GB per GPU)
Network cardsConnectX-7, 400GConnectX-8, 800G
CoolingDirect liquid coolingDirect liquid cooling
AvailabilityShipping since 2024-2025Ramping through 2026

Beyond GB300, NVIDIA’s next architecture is Vera Rubin, expected in the second half of 2026. The Vera Rubin NVL144 generation is expected to bring Vera CPUs, HBM4 memory, and higher per-GPU power. For buyers, the practical implication is that each GB200-generation rack you deploy today should use optics and cabling that carry forward to GB300, where 800G and 1.6T OSFP links become standard.

Frequently Asked Questions

What is the NVIDIA GB200 superchip?

The NVIDIA GB200 superchip is a Blackwell-based accelerated computing module that combines one Grace CPU with two Blackwell B200 Tensor Core GPUs.

It is the foundation of NVIDIA GB200 NVL72 and NVL36x2 rack-scale AI systems designed for large-scale AI training and inference workloads.

What is the difference between GB200 NVL72 and NVL36x2?

The GB200 NVL72 integrates 72 Blackwell GPUs and 36 Grace CPUs inside a single liquid-cooled rack.

The GB200 NVL36x2 distributes the same GPU scale across two racks, providing more flexible deployment options for facilities with different power and cooling constraints.

Why does GB200 require liquid cooling?

GB200 systems achieve extremely high compute density, resulting in rack power levels beyond traditional air-cooled infrastructure.

Direct-to-chip liquid cooling removes heat more efficiently from high-power components such as GPUs, CPUs, and NVLink switch ASICs.

What optics and cables does a GB200 rack use?

GB200 networking depends on the specific deployment architecture.

Typical connectivity includes:

  • 400G QSFP-DD optical modules for current-generation AI fabrics
  • 800G OSFP optical modules for next-generation AI networking
  • High-performance DAC solutions for short-distance connections

The exact module selection depends on the switch platform, protocol, and link distance.

Is GB200 still relevant now that GB300 is shipping?

Yes.

GB200 remains an important AI infrastructure platform because it introduced NVIDIA’s rack-scale Blackwell architecture and established the design foundation for future generations.

GB300 expands this architecture with higher performance and increased adoption of 800G networking.

Conclusion

The NVIDIA GB200 superchip represents a major transformation in AI infrastructure design.

Instead of scaling performance through individual GPU servers, GB200 introduces rack-scale computing platforms that combine:

  • High-density GPU acceleration
  • NVLink scale-up networking
  • Liquid cooling
  • Advanced optical connectivity

The key takeaways are:

  • GB200 combines one Grace CPU with two Blackwell B200 GPUs to create a high-performance AI compute module.
  • GB200 NVL72 integrates 72 GPUs and 36 CPUs into a single liquid-cooled rack.
  • GB200 NVL36x2 provides a flexible two-rack deployment model while maintaining a large NVLink domain.
  • AI cluster networking is moving from 400G toward 800G optical connectivity.
  • Choosing scalable optical solutions is critical as AI infrastructure evolves toward GB300 and future platforms.

For organizations building NVIDIA GB200 or GB300 AI clusters, selecting the right optical modules, DAC cables, and networking components is essential for achieving reliable high-performance connectivity.

FiberMall provides NVIDIA-compatible optical transceiver solutions, including 400G and 800G OSFP modules, designed for next-generation AI data center networks.

Scroll to Top