Marcus stared at the link-down alarm on his console. His team had just swapped a 200G HDR switch for a 400G NDR switch.
The cables were right. The transceivers seated cleanly. The ports showed green on both sides.
Yet the fabric wouldn’t come up.
He called the vendor’s support line. The first question surprised him: “Did you migrate your Subnet Manager configuration?” Marcus had spent years managing Ethernet networks. He knew VLANs, STP, and BGP. He hadn’t realized that InfiniBand architecture hands control of addressing and routing to a separate piece of software entirely.
That’s the gap this guide fills. InfiniBand isn’t just faster Ethernet. It is a purpose-built, lossless, switched-fabric interconnect with its own layers, components, and management model. Once you understand how those pieces fit together, decisions about switches, adapters, cables, and topologies become much clearer.
In this article, we’ll walk through InfiniBand architecture from the physical port to the application layer. You’ll learn what Host Channel Adapters, switches, and Subnet Managers actually do, how the five layers carry traffic, and why 2026 is a turning point for NDR, XDR, and GDR fabrics.
Want the cabling side of the story first? Read our InfiniBand cables guide for a practical look at DAC, AOC, and connector choices.

Table of Contents
ToggleInfiniBand Architecture at a Glance
InfiniBand architecture is a high-speed, low-latency, switched-fabric interconnect standard designed for high-performance computing, AI training clusters, and low-latency storage. Every device connects through dedicated point-to-point links to switches, eliminating the shared-medium collision domains you see in classic Ethernet.
The architecture separates control from data. A central Subnet Manager discovers the fabric, assigns addresses, and programs forwarding tables into switches. Data packets then move through the fabric at line rate with hardware-offloaded RDMA. This separation is what makes InfiniBand predictable, but it also means the architecture has more moving parts than a typical IP network.
At the highest level, the fabric consists of five building blocks:
Host Channel Adapter (HCA), the network adapter that terminates InfiniBand in a server, GPU, or storage controller.
Switch, forwards packets between ports based on Local IDs (LIDs) programmed by the Subnet Manager.
Router, connects separate InfiniBand subnets using Global IDs (GIDs).
Subnet Manager (SM), the control-plane brain that configures and monitors the fabric.
Physical links, the cables, connectors, and transceivers that carry signaling.
If you’re comparing this stack to Ethernet, start with our InfiniBand vs Ethernet overview.
Physical Layer: Ports, Links, and Generations
The physical layer defines how bits move across the wire, fiber, or cable. InfiniBand generations scale by adding lanes and increasing signaling rate, but the architecture itself stays consistent.
| Generation | Per-Lane Rate | Lanes | Aggregate Bandwidth | Common Connector |
| SDR | 2.5 Gb/s | 4 | 10 Gb/s | SDR CX4 |
| DDR | 5 Gb/s | 4 | 20 Gb/s | QSFP-like |
| QDR | 10 Gb/s | 4 | 40 Gb/s | QSFP |
| FDR | 14 Gb/s | 4 | 56 Gb/s | QSFP |
| EDR | 25 Gb/s | 4 | 100 Gb/s | QSFP28 |
| HDR | 50 Gb/s | 4 | 200 Gb/s | QSFP56 |
| NDR | 100 Gb/s | 4 or 8 | 400 Gb/s | OSFP / QSFP112 |
| XDR | 100 Gb/s | 8 | 800 Gb/s | OSFP |
| GDR | 200 Gb/s | 8 | 1.6 Tb/s | OSFP-XD / CPO |
The encoding has evolved too. Early generations used 8b/10b. HDR moved to 64b/66b. NDR and XDR rely on PAM4 signaling with forward error correction to push 400G and 800G through the same form-factor envelopes.
This is where physical procurement meets architecture. You can’t drop an NDR HCA into an HDR switch and expect it to negotiate down like Ethernet sometimes does. The protocol supports backward negotiation, but the connector and firmware qualification have to match. FiberMall tests its InfiniBand cables and modules across these generations to reduce that risk.
Need a refresher on speeds and encoding? See our InfiniBand generations speed chart.
Link Layer: Lossless Fabric and Credit-Based Flow Control
The link layer is where InfiniBand starts to feel different from Ethernet. Instead of best-effort forwarding with drop-and-recover behavior, InfiniBand uses credit-based flow control. A sender can only transmit when the receiver has advertised enough buffer credits. Packets don’t drop due to congestion inside a well-managed fabric.

Each packet carries a Local Route Header (LRH) with a destination Local ID (LID). Switches read that LID and forward the packet using tables the Subnet Manager installed.
No MAC learning. No broadcast ARP. No spanning tree.
The fabric is deterministic because the control plane pre-computes paths.
The link layer also handles packet framing, error detection, and link-level retry. If a link is noisy, the receiver can request a retransmission before the packet ever reaches the transport layer. That matters for RDMA, where the application expects direct memory updates without TCP’s retry machinery.
Sarah learned this the hard way. Her team ran a large AI cluster on a single InfiniBand subnet. One day, a misconfigured maintenance window rebooted the primary Subnet Manager. The standby SM took over, but its cached LID map was stale.
For about ninety seconds, some switches forwarded packets to ports that no longer matched the active LID assignments. The fabric didn’t drop packets randomly; it delivered them to the wrong destinations.
Training jobs crashed with silent data corruption rather than obvious timeouts.
The fix was simple in hindsight: always run redundant Subnet Managers with synchronized state. The lesson was deeper. In InfiniBand, the link layer depends entirely on the control plane being consistent.
Network Layer: Subnets, LIDs, GIDs, and Routers
The network layer handles addressing and routing. Inside a subnet, every port has a 16-bit Local ID (LID). The Subnet Manager assigns LIDs during fabric initialization and can reassign them when topology changes. The link-layer LRH uses the LID to route packets through switches.
When traffic needs to cross subnet boundaries, the network layer uses Global IDs (GIDs). A GID is a 128-bit identifier derived from the port’s GUID and a subnet prefix. InfiniBand routers read the Global Route Header (GRH) and forward packets between subnets.
Most AI and HPC clusters today live in a single subnet. Routers become relevant when connecting multiple clusters, bridging to a management network, or extending InfiniBand across physically separate data halls.
The Subnet Manager deserves more attention here because it is the architectural control point. It runs as software on a server, switch, or dedicated appliance. It discovers the fabric through Subnet Manager Agents (SMAs) embedded in every device. Then it calculates routing tables, configures partitions, sets QoS, and monitors link status.
Without it, the fabric cannot initialize.
This centralized model is powerful. It eliminates broadcast storms and enables fast reconvergence. But it also creates a single point of failure unless you deploy a backup SM. Modern NVIDIA fabrics typically run the SM on one or more management nodes with automatic failover.
Transport Layer: Queue Pairs, RDMA, and Kernel Bypass
The transport layer is where applications interact with the fabric. InfiniBand does not expose a socket API by default. Instead, applications create Queue Pairs (QPs). Each QP is a pair of work queues: one for sends and one for receives.
Completion Queues (CQs) notify the application when operations finish.
There are two main transport modes:
Send/Receive, the sender posts a buffer, the receiver posts a buffer, and the HCA matches them.
RDMA Read/Write, one HCA directly reads from or writes to memory on a remote node without involving the remote CPU.
RDMA is the headline feature. It lets a GPU on one server pull data directly from GPU memory on another server. The CPU and operating system are bypassed entirely. That’s why InfiniBand latency is measured in microseconds while TCP/IP over Ethernet often sits in tens of microseconds.

Kernel bypass also matters. In traditional networking, every packet crosses into kernel space, then into user space, then back. InfiniBand Verbs and RDMA libraries let applications post work directly to the HCA from user space. The HCA handles segmentation, reassembly, reliability, and ordering in hardware.
For AI training frameworks like NCCL and MPI implementations, this means all-reduce operations can move huge tensors with minimal CPU overhead. The transport layer doesn’t just carry data; it offloads the network from the compute.
Upper-Layer Protocols: MPI, NCCL, IPoIB, and Storage
Above the transport layer sit the protocols that applications actually use. The most important ones in AI/HPC are MPI and NCCL. They use InfiniBand Verbs to implement collectives like all-reduce, broadcast, and gather. NVIDIA’s NCCL is tuned specifically for GPU-to-GPU communication over InfiniBand and NVLink.
IPoIB lets TCP/IP traffic run over InfiniBand. It is useful for management, file transfers, or applications that don’t speak native InfiniBand Verbs. But it doesn’t deliver the latency benefits of RDMA, because it still passes through the IP stack.
Storage protocols include SRP (SCSI RDMA Protocol) and NVMe-oF. These let block storage traffic bypass the CPU and move directly between initiator and target memory. In high-performance storage arrays, this is often the reason InfiniBand is chosen over Fibre Channel or iSCSI.
These upper-layer protocols are why the architecture matters in practice. The same physical and transport layers support AI training, parallel file systems, and IP workloads. The choice of protocol depends on the application, not the cabling.
InfiniBand Architecture vs Ethernet
Engineers often ask whether InfiniBand is worth the complexity when Ethernet keeps getting faster. The honest answer is: it depends on the workload.
| Aspect | InfiniBand | Ethernet |
| Fabric model | Switched fabric with centralized SM | Broadcast/multi-access with distributed control |
| Flow control | Credit-based, lossless by design | Best-effort; lossless requires PFC, ECN, tuning |
| Latency | Sub-microsecond, deterministic | Higher, more variable |
| CPU offload | Hardware RDMA and kernel bypass | Software TCP/IP unless using RoCE/iWARP |
| Routing | LID-based, pre-computed by SM | MAC/IP learning, BGP, OSPF |
| Management | Subnet Manager centralizes config | Switch-by-switch or controller-based |
| Ecosystem | NVIDIA-dominated (Mellanox heritage) | Multi-vendor, open |
Ethernet has closed much of the gap with RoCEv2 and the Ultra Ethernet Consortium. But RoCE still relies on PFC and ECN to approximate lossless behavior, and it requires careful congestion management. InfiniBand gives you lossless behavior out of the box because the architecture was designed for it.
Ethernet wins in mixed-workload clouds and environments where interoperability across many vendors matters. InfiniBand wins when a single owner wants the lowest possible latency and the simplest path to RDMA at scale.

2026 Roadmap: NDR, XDR, and GDR Architectures
Right now, NDR 400G is the production standard for H100 and H200 GPU clusters. NVIDIA Quantum-2 switches with ConnectX-7 adapters are the common building blocks. XDR 800G is shipping with Quantum-X800 and ConnectX-8 for frontier AI builds.
GDR at 1.6 Tb/s is on the roadmap for late 2026 or early 2027. The architecture won’t change conceptually, but the physical layer will. Higher lane rates and more lanes mean more power per port.
An XDR optic can draw 20–30 W. A fully populated 144-port Quantum-X800 switch can add more than 2,000 W of transceiver heat alone.
Form factors are shifting too. Switches are standardizing on OSFP for thermal headroom. NICs and DPUs often stay with QSFP112 because of size constraints. That means breakout cables will remain central to mixed-generation architectures.
For a forward-looking deployment view, read our 800G NDR InfiniBand deployment guide.
InfiniBand Architecture Checklist for Buyers
If you’re planning a deployment, use this checklist before you order hardware:
1. Generation and speed, HDR, NDR, NDR200, XDR, or GDR?
2. Port form factor, QSFP56, OSFP, or QSFP112?
3. Subnet Manager placement, where will the primary and standby SM run?
4. Topology, fat-tree, rail-optimized, or dragonfly+?
5. Cable type, passive DAC, active copper, AOC, or pluggable optics?
6. Power and cooling, have you budgeted transceiver heat at 800G and above?
7. Firmware and EEPROM, are cables and modules qualified for your exact switch/NIC revision?
FiberMall tests its InfiniBand cables and optical modules against major NVIDIA platforms. If you’re unsure which generation or form factor fits your architecture, explore our InfiniBand collection and request a compatibility check.
Conclusion
InfiniBand architecture is more than fast ports and low latency. It is a complete stack: physical links, a lossless link layer, a centrally managed network layer, a hardware-offloaded transport layer, and upper-layer protocols that feed directly into AI, HPC, and storage workloads.
Start with the physical layer. Match generation, connector, and cable to the hardware you already own. Then make sure your Subnet Manager is redundant and your topology is well understood. After that, the transport and application layers mostly take care of themselves.
Get the architecture right and you’ll avoid Marcus’s link-down surprise and Sarah’s silent fabric partition. Get it wrong and even the most expensive switch becomes a very expensive troubleshooting exercise.
If you’re planning an InfiniBand deployment, contact FiberMall for a compatibility review or a quote on tested InfiniBand cables, optical modules, and breakout solutions.
Related Products:
-
NVIDIA MMA4Z00-NS400 Compatible 400G OSFP SR4 Flat Top PAM4 850nm 30m on OM3/50m on OM4 MTP/MPO-12 Multimode FEC Optical Transceiver Module
$400.00
-
NVIDIA MMS4X00-NS400 Compatible 400G OSFP DR4 Flat Top PAM4 1310nm MTP/MPO-12 500m SMF FEC Optical Transceiver Module
$450.00
-
NVIDIA MMA1Z00-NS400 Compatible 400G QSFP112 VR4 PAM4 850nm 50m MTP/MPO-12 OM4 FEC Optical Transceiver Module
$385.00
-
NVIDIA MMS1X00-NS400 Compatible 400G NDR QSFP112 DR4 PAM4 1310nm 500m MPO-12 with FEC Optical Transceiver Module
$500.00
-
NVIDIA MCP7Y70-H001 Compatible 1m (3ft) 400G Twin-port 2x200G OSFP to 4x100G QSFP56 Passive Breakout Direct Attach Copper Cable
$120.00
-
NVIDIA MCP7Y60-H001 Compatible 1m (3ft) 400G OSFP to 2x200G QSFP56 Passive Direct Attach Cable
$99.00
-
NVIDIA MFA7U10-H003 Compatible 3m (10ft) 400G OSFP to 2x200G QSFP56 twin port HDR Breakout Active Optical Cable
$750.00
-
NVIDIA MMA4Z00-NS Compatible 800GBASE 2 x SR4/SR8 OSFP PAM4 850nm 100m DOM Dual MPO-12 MMF Optical Transceiver Module
$550.00
-
NVIDIA MMA4Z00-NS-FLT Compatible 800GBASE 2 x SR4/SR8 OSFP RHS/Flat Top PAM4 850nm 100m DOM Dual MPO-12 MMF Optical Transceiver Module
$600.00
-
NVIDIA MMS4X00-NM Compatible 800GBASE 2 x DR4/DR8 OSFP IHS/Closed Finned Top PAM4 1310nm 500m DOM Dual MTP/MPO-12 SMF Optical Transceiver Module
$600.00
-
NVIDIA MMS4X00-NM-FLT Compatible 800GBASE 2 x DR4/DR8 OSFP Flat Top PAM4 1310nm 500m DOM Dual MTP/MPO-12 SMF Optical Transceiver Module
$650.00
-
NVIDIA MMS4X50-NM Compatible 800G 2x FR4 OSFP IHS/Closed Finned Top PAM4 1310nm 2km DOM Dual Duplex LC SMF InfiniBand NDR Optical Transceiver Module
$1000.00
-
NVIDIA MCP7Y00-N001 Compatible 1m (3ft) 800Gb Twin-port OSFP to 2x400G OSFP InfiniBand NDR Breakout Direct Attach Copper Cable
$160.00
-
NVIDIA MCA7J60-N004 Compatible 4m (13ft) 800G Twin-port OSFP to 2x400G OSFP InfiniBand NDR Breakout Active Copper Cable
$800.00
-
NVIDIA MMS4A00 (980-9IAH1-00XM00) Compatible 1.6T 2 x DR4/DR8 OSFP224 PAM4 1311nm 500m IHS/Finned Top Dual MPO-12 SMF Optical Transceiver Module
$1500.00
-
NVIDIA MMS4A50 Compatible 1.6T 2xFR4/FR8 OSFP224 PAM4 1310nm 2km IHS/Finned Top Dual Duplex LC SMF Optical Transceiver Module
$1800.00
-
NVIDIA MMS4A00-RHS Compatible 1.6T 2xDR4/DR8 OSFP224 PAM4 1311nm 500m RHS/Flat Top Dual MPO-12/APC InfiniBand XDR SMF Optical Transceiver Module
$2000.00
-
NVIDIA MCA7K20-X001 Compatible 1m (3ft) Twin-port 2x800Gb/s OSFP224 IHS/Finned Top to 4x400Gb/s OSFP224 RHS/Flat Top InfiniBand XDR Active Copper Splitter Cable
$2059.00
Related Posts
- What Is InfiniBand? The Complete Guide to High-Performance AI Networking
- InfiniBand vs RoCEv2: AI Data Center Networking Guide (2026)
- InfiniBand vs Ethernet: Which Network Should Your AI Cluster Use?
- InfiniBand Generations: SDR to GDR Speed Chart (2026)
- InfiniBand Cables: DAC vs AOC & Connector Guide (2026)
- InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale
- NVIDIA InfiniBand Switches: Quantum-2, Quantum-X800, and 2026 Roadmap
- ConnectX-7 vs ConnectX-8: Which NVIDIA SuperNIC Fits Your AI Cluster?
- InfiniBand Subnet Manager: OpenSM, Switch SM & UFM 2026
- InfiniBand Price Guide 2026: Switches, HCAs, Cables & TCO
- InfiniBand Troubleshooting: A Step-by-Step Guide for 2026
- 800G NDR InfiniBand Deployment Guide: Quantum-2 & ConnectX-7
