NVIDIA InfiniBand Switches: Quantum-2, Quantum-X800, and 2026 Roadmap

The rack was empty except for the PDUs, rack rails, and several boxes of cables. Marcus, the lead infrastructure engineer, had just unboxed two NVIDIA Quantum-2 switches for a new AI training cluster. His team had planned the network topology for weeks. But when they started connecting DACs, something was wrong. The cables would not fully seat in the switch ports. They had ordered QSFP112-to-QSFP112 direct-attach copper cables for switches that required OSFP on the switch side. The mistake cost them fourteen days and a rushed optics order that blew the cable budget.

Choosing an InfiniBand switch is not just about port count and bandwidth. Connector form factors, power delivery, thermal design, optics compatibility, and cable availability all determine whether your fabric comes up smoothly or stalls before the first training job runs.

This guide covers NVIDIA’s InfiniBand switch lineup from Quantum-2 through Quantum-X800 and the 2026 roadmap. You will learn how InfiniBand switches differ from Ethernet switches, what specs matter, which cables and optics you need, and how to plan power and procurement. Whether you are building a 64-GPU lab or a hyperscale training cluster, this article will help you spec the right switch.

New to InfiniBand? Read the complete InfiniBand guide first โ†’

The rack was empty except for the PDU

What Is an InfiniBand Switch?

An InfiniBand switch is a specialized network switch designed for low-latency, lossless communication in high-performance computing and AI clusters. Unlike traditional Ethernet switching, which commonly relies on MAC address learning and Ethernet control-plane protocols, an InfiniBand switch forwards packets based on a Local Identifier, or LID, assigned by a central Subnet Manager.

The switch itself is simple in concept. It receives a packet, looks up the destination LID in a local forwarding table, and sends the packet out the correct port. There is no MAC learning in the Ethernet sense. There is no need for a distributed Layer 2 learning process across the fabric. The intelligence lives in the Subnet Manager, which discovers the fabric and programs forwarding tables across the switches.

This design makes InfiniBand switches extremely fast and predictable. It also means that if the Subnet Manager is not running or is misconfigured, the fabric may fail to initialize correctly, and topology changes may not be programmed as expected. Most production deployments run a primary Subnet Manager on a management node, with a standby instance for redundancy.

How InfiniBand Switching Works

InfiniBand forwarding is cut-through and non-blocking in modern high-performance switches. When a packet arrives, the switch reads the destination LID and forwards the packet with minimal delay. It does not need to buffer the entire packet before deciding where to send it. That helps keep switch latency in the sub-microsecond range.

The Subnet Manager discovers all switches and host channel adapters in the fabric. It assigns each port a LID, calculates forwarding paths, and downloads forwarding tables to each switch. It also configures service levels, virtual lanes, and other fabric parameters that allow different traffic classes to share the physical network while reducing head-of-line blocking.

A key feature in NVIDIA InfiniBand switches is SHARP, or Scalable Hierarchical Aggregation and Reduction Protocol. SHARP offloads supported collective operations into the switch network. Instead of forcing every reduction operation to be handled entirely by endpoints, the fabric can perform part of the aggregation inside the network. This reduces endpoint overhead and can improve scaling efficiency for MPI, HPC, and AI training workloads.

InfiniBand Switch

NVIDIA InfiniBand Switch Lineup

NVIDIA’s InfiniBand switch family has two main deployment tiers in 2026, plus a forward-looking roadmap tier. Understanding the differences is essential for matching hardware to workload, topology, power budget, and cable plan.

Quantum-2 (QM9700 / QM9790)

Quantum-2 is the current workhorse for NDR 400G InfiniBand. A single Quantum-2 switch provides 64 ports of 400G, giving 51.2 terabits per second of aggregate switching capacity. It uses OSFP form-factor ports and supports SHARPv4.

There are two main variants:

QM9700: Air-cooled, standard for most data centers.

QM9790: Liquid-cooled, designed for high-density deployments where traditional airflow is insufficient.

Quantum-2 is the switch you will find in many AI training clusters built between 2023 and 2026. It commonly pairs with NVIDIA ConnectX-7 host channel adapters and BlueField-3 DPUs.

Quantum-X800

Quantum-X800 moves to XDR 800G InfiniBand. It doubles the per-port bandwidth compared with Quantum-2 and is designed for very large AI and HPC fabrics.

The key point is that Quantum-X800 is not a single fixed 64-port switch. The family includes different chassis configurations. The Q3200 is a 2U system with two independent 28.8 Tb/s switches, each providing 36 ports of 800G. The Q3400 is a 4U system that provides 144 ports of 800G distributed across 72 OSFP cages, with 115.2 Tb/s of switching throughput.

Quantum-X800 introduces SHARPv4, enhanced adaptive routing, telemetry-based congestion control, and improved power-efficiency features. It is designed to pair with NVIDIA ConnectX-8 SuperNICs and NVIDIA LinkX XDR cables and transceivers. For new high-end AI cluster builds in 2026, Quantum-X800 is the platform to consider when 800G end-to-end bandwidth is required.

Quantum-2 (QM9700 QM9790)

Quantum-X1600 Roadmap

Looking beyond XDR, the industry roadmap is moving toward 1.6T-class InfiniBand connectivity in the GDR era. NVIDIA has already shown the direction of travel with Quantum-X800, ConnectX SuperNIC roadmap messaging, and silicon-photonics-based Quantum-X designs.

For practical 2026 procurement, however, Quantum-2 and Quantum-X800 are the relevant platforms to specify. 1.6T-class InfiniBand should be treated as roadmap planning rather than a standard deployment option. Buyers should avoid locking procurement language around a future switch generation until NVIDIA publishes final product specifications, connector details, cable options, power requirements, and availability.

InfiniBand Switch Specs Comparison

SpecificationQuantum-2 QM9700Quantum-X800
GenerationNDR 400GXDR 800G
Ports64 x 400G64 x 800G
Aggregate Bandwidth51.2 Tbps~102.4 Tbps
Port Form FactorOSFPOSFP
Switching LatencySub-microsecondSub-microsecond
SHARPv4v4+
CoolingAir (QM9700) / Liquid (QM9790)Air / Liquid options
Typical Power~800-1,200W~1,200-1,800W
ManagementSubnet Manager / UFMSubnet Manager / UFM

The numbers are impressive, but the practical takeaway is this: Quantum-2 is the proven 400G NDR platform with a mature 400G cable and optics ecosystem. Quantum-X800 is the 800G XDR platform for new high-end builds. Your choice depends on whether you need maximum per-port bandwidth now, whether your hosts are ready for 800G, and whether your rack power and cooling design can support the newer switch generation.

Port Form Factors: OSFP vs QSFP112

This is where procurement mistakes happen. Quantum-2 switches use OSFP cages. OSFP is an eight-lane form factor. It is larger than QSFP and is not mechanically compatible with QSFP-family connectors.

QSFP112 is a four-lane form factor often used on NVIDIA ConnectX-7 host channel adapters for NDR 400G. A ConnectX-7 adapter may have QSFP112 ports, while the Quantum-2 switch it connects to has OSFP cages. That means the cable between them must have the correct connector on each end: OSFP on the switch side and QSFP112 on the host side.

Ordering 200 QSFP112-to-QSFP112 DACs for a Quantum-2 rollout is exactly the kind of error that delays clusters. Always confirm the port form factor on both ends of every link before ordering cables.

For Quantum-X800, the same rule applies. Do not assume that an 800G cable will fit just because the speed is correct. Confirm the exact switch model, cage type, cable connector, cable length, airflow direction, and firmware support before placing a production order.

Port Form Factors OSFP vs QSFP112

Cabling an InfiniBand Switch

InfiniBand switches support three main cable types:

DAC (Direct Attach Copper): Lowest cost, shortest reach, typically 1-3 meters. Passive DAC draws almost no power but has strict distance limits. Active DAC extends reach slightly.

AOC (Active Optical Cable): Pre-terminated optical cable with transceivers built in. Common for 3-30 meter links. Easier to deploy than separate transceivers and fiber.

Optical transceivers plus fiber: Most flexible for longer distances and structured cabling. Requires matching form factor, wavelength, and fiber type.

For NDR 400G, OSFP ports accept OSFP modules or breakout cables. A 400G OSFP port can break out to 2x200G QSFP56 or 4x100G QSFP28 using the right cable. Breakouts are essential for fat-tree spine-leaf designs where leaf switches run at lower speeds than spine switches.

Power, Thermal, and Rack Planning

High-speed InfiniBand switches consume serious power. A fully loaded Quantum-2 air-cooled switch typically draws 800 to 1,200 watts. Quantum-X800 pushes past 1,200 watts and can approach 1,800 watts depending on optics and traffic.

That power becomes heat. A single 42U rack with four Quantum-2 switches and fully populated optics can add over 5,000 watts of heat load. One data center team we know planned for three switches per rack and discovered their cooling could only handle two. They had to redistribute hardware across additional racks and rethink cable management.

Before racking, verify:

Rack power distribution capacity per PDU

Cooling capacity for the row or room

Airflow direction (front-to-back or reverse)

Cable bend radius and management

Weight load per rack

Power, Thermal, and Rack Planning

Cost and Procurement

NVIDIA InfiniBand switches are expensive. A new Quantum-2 switch can cost tens of thousands of dollars, and optics or active cables can add a significant amount per port. During peak AI build-outs, lead times can stretch from weeks to months depending on the model, region, and supply channel.

For smaller labs and research groups, the used market offers alternatives. Older Mellanox switches such as the SX6036 provide FDR 56G InfiniBand at a fraction of the cost. They are not suitable for bleeding-edge AI training, but they are useful for learning, smaller HPC clusters, and development environments. For used EDR 100G InfiniBand, buyers should look at newer EDR-generation Mellanox/NVIDIA platforms rather than SX6036.

Third-party MSA-compliant optics and cables can reduce cabling costs compared with NVIDIA-branded LinkX parts. The key is qualification. Test sample modules in your actual switch firmware and HCA firmware before placing a large order. FiberMall supplies InfiniBand-compatible transceivers and cables tested for NVIDIA Quantum-2 and ConnectX-7 deployments.

A university lab we worked with reduced its InfiniBand cable budget by qualifying third-party OSFP AOCs before rollout. The team tested sample cables under sustained traffic, monitored bit error rates and temperatures, and then placed the full order. The savings helped fund additional compute capacity.

Deployment Checklist

Before you put production traffic on an InfiniBand switch, run through this checklist:

1. Verify form factors: Confirm OSFP or QSFP112 on every port end.

2. Check firmware: Switch, HCA, and Subnet Manager firmware should be compatible.

3. Qualify cables: Test DAC, AOC, or optical modules for bit errors.

4. Deploy Subnet Manager: Primary and standby instances on management nodes.

5. Configure virtual lanes and QoS: Match your workload requirements.

6. Run fabric diagnostics: Use ibdiagnet or equivalent tools.

7. Benchmark bandwidth: Run ib_write_bw or MPI/NCCL tests.

8. Burn in: Run sustained traffic for 24-72 hours before production.

9. Monitor thermals: Watch switch and optic temperatures during burn-in.

10. Document cabling: Label every cable clearly. You will thank yourself later.

FAQ

Which switches are used in InfiniBand networks?
Modern InfiniBand networks use specialized low-latency switches from NVIDIA. The main current platforms are Quantum-2 for 400G NDR and Quantum-X800 for 800G XDR.

How many ports does Quantum-2 have?
A Quantum-2 switch provides 64 ports of 400G NDR InfiniBand or up to 128 ports of 200G through port splitting, using 32 physical OSFP cages.

Can I use third-party optics in InfiniBand switches?
Yes, if they are MSA-compliant and qualified in your specific switch and HCA firmware. Always test samples before large orders.

What is the difference between Quantum-2 and Quantum-X800?
Quantum-2 supports 400G NDR per port. Quantum-X800 supports 800G XDR per port with roughly double the aggregate switching capacity.

How much power does an InfiniBand switch draw?
A fully loaded Quantum-2 typically draws 800-1,200 watts. Quantum-X800 can draw 1,200-1,800 watts depending on configuration.

What cable do I need for a Quantum-2 switch?
Quantum-2 uses OSFP ports. Choose OSFP DAC, AOC, or optical transceivers. If connecting to a QSFP112 host adapter, use a cable with the correct ends on each side.

Is InfiniBand switching lossless?
Yes. InfiniBand uses credit-based flow control, so packets are not dropped due to congestion.

How is an InfiniBand switch managed?
InfiniBand switches are managed centrally by a Subnet Manager, which programs forwarding tables and fabric configuration. Operators do not configure each switch individually like Ethernet.

Conclusion

NVIDIA InfiniBand switches are the backbone of modern AI training fabrics. Quantum-2 delivers proven 400G performance for most clusters. Quantum-X800 doubles that to 800G for the most demanding workloads. The Quantum-X1600 roadmap promises 1.6 Tbps in the near future.

The biggest mistakes in InfiniBand switch deployments are not about the switch silicon. They are about form factors, cables, power, and thermal planning. Get the OSFP versus QSFP112 question right. Size your racks for real power draw. Qualify your optics. And never skip the burn-in.

A well-planned InfiniBand fabric is invisible to the workloads running on it. GPUs stay fed, gradients synchronize quickly, and training runs finish on time. If you need help sourcing MSA-compliant InfiniBand cables and transceivers tested for NVIDIA Quantum switches, FiberMall can support your specification and procurement.

Scroll to Top