A single 1% packet-drop rate on a 1,024-GPU H100 cluster can burn roughly $250,000 per week in idle compute. That isn’t a theory. It’s why hyperscalers and HPC labs obsess over lossless, low-latency interconnects. InfiniBand RDMA is the technology that makes those interconnects possible.
Remote Direct Memory Access (RDMA) lets one server read from or write to another server’s memory without waking either CPU or copying data through the operating system kernel. InfiniBand was built specifically for this. While Ethernet with RoCEv2 can run RDMA over a familiar stack, InfiniBand RDMA offers native losslessness and sub-microsecond latency that remains hard to match.
In this article, we’ll walk through how InfiniBand RDMA actually works: verbs, queue pairs, memory regions, and the connection setup flow. Then we’ll connect it to AI training, GPUDirect RDMA, and the physical layer that often gets underestimated.
Table of Contents
ToggleWhat Is InfiniBand RDMA?
InfiniBand RDMA is a networking technology that moves data directly from one application’s memory to another application’s memory across a switched fabric. The CPU doesn’t copy buffers. The kernel doesn’t process packets. The Host Channel Adapter (HCA), such as NVIDIA ConnectX-7, handles segmentation, reliability, flow control, and ordering in silicon.

Traditional TCP/IP follows a longer path. Data copies from user space to kernel socket buffers, then to the NIC, then across the wire. The receiver reverses every step. Each copy and context switch adds latency and consumes CPU cycles. RDMA removes most of that overhead.
The result is three core advantages:
Kernel bypass: Applications post work requests directly to the NIC through user-space libraries.
Zero-copy: Data stays in pinned user buffers; it never transits kernel socket buffers.
NIC offload: Transport logic runs on the HCA, not the host CPU.
For AI clusters, this matters because training workloads generate enormous east-west traffic. A single all-reduce step across a few hundred GPUs can move terabytes of gradients. InfiniBand RDMA keeps that traffic off the CPU and out of the kernel.
These three properties define InfiniBand RDMA and separate it from conventional networking.
How InfiniBand RDMA Works: The Software Stack
Most competitors stop at “RDMA bypasses the kernel.” That is true, but it skips the part that actually determines whether your code works. InfiniBand RDMA has a distinct software model built around verbs, queue pairs, and memory regions.
Verbs: The RDMA API
Verbs are the operations an application uses to talk to the HCA. They are defined by the InfiniBand Architecture specification and implemented in libraries like libibverbs and rdma-core. A verb describes an action: post a send request, create a queue pair, register memory, modify a queue-pair state.
The verbs API is lower level than sockets. There is no implicit buffering and no standard blocking send(). The application manages memory, queues, and completions explicitly. This is the interface user-space code uses to drive InfiniBand RDMA.
Queue Pairs: The Communication Endpoint
A Queue Pair (QP) is the core communication unit in InfiniBand RDMA. Every QP contains a Send Queue (SQ) and a Receive Queue (RQ). The application posts Work Requests (WRs) to these queues. The HCA consumes them and, when finished, writes Work Completions (WCs) to a Completion Queue (CQ).
QPs support different transport services. Reliable Connection (RC) is the most common: it guarantees in-order delivery and uses acknowledgments. Unreliable Datagram (UD) is lighter and useful for some MPI patterns. Shared Receive Queues let multiple QPs share one RQ to save memory.

Memory Regions: Pinned and Keyed
Before the HCA can touch application memory, that memory must be registered as a Memory Region (MR). Registration pins the pages so they cannot be swapped out and builds a virtual-to-physical mapping the NIC can use.
Each MR gets two keys: an L_Key for local access and an R_Key for remote access. For RDMA READ or WRITE, the initiator must know the remote virtual address and the remote R_Key. That information is typically exchanged over a TCP socket or similar out-of-band channel before RDMA traffic starts.
Memory registration has overhead. Pinning large buffers takes time and consumes kernel resources. Some newer systems use On-Demand Paging (ODP) to avoid explicit registration, but ODP can introduce latency spikes if page mappings are not ready.
Protection Domains and Completion Queues
A Protection Domain (PD) groups QPs and MRs that belong to the same address-space context. QPs and MRs must share a PD to interact. Completion Queues collect notifications from multiple QPs, and the application can poll them or wait for events.
Polling is faster but burns CPU. Event-driven notification saves CPU but adds latency. Most high-performance workloads poll.
Connection Setup Flow
A QP starts in RESET. The application moves it through INIT, then RTR (Ready to Receive), then RTS (Ready to Send). Only in RTS can both sides initiate sends. Each transition requires correct parameters: Local Identifier (LID), Queue Pair Number (QPN), Packet Sequence Number (PSN), and path MTU. Get one wrong and the QP throws a syndrome error that can take time to trace.
RDMA Operations: SEND/RECEIVE, READ, WRITE
InfiniBand RDMA supports several transfer semantics. Choosing the right one affects both performance and complexity.
| Operation | Sidedness | Remote CPU Involved? | Best For |
| SEND / RECEIVE | Two-sided | Yes, receiver posts buffers | MPI messages, dynamic communication |
| RDMA READ | One-sided | No, initiator pulls data | Metadata lookups, polling reads |
| RDMA WRITE | One-sided | No, initiator pushes data | Bulk transfers, checkpointing |
| Atomic | One-sided | No, hardware does RMW | Locks, counters |
SEND/RECEIVE is conceptually closest to messaging. The receiver must pre-post receive buffers. If a SEND arrives with no posted receive, the QP typically enters an error state.
RDMA READ and WRITE are true one-sided operations. The remote CPU does not participate. For a WRITE, the initiator supplies a remote address and R_Key; the remote HCA DMAs data directly into that memory. This is where GPUDirect RDMA becomes possible.
Picking the right operation is a central part of InfiniBand RDMA tuning.
Why InfiniBand RDMA Dominates AI Training
AI training is unusually sensitive to network behavior. Distributed SGD requires frequent gradient synchronization. A delayed packet at one GPU can stall the entire all-reduce, leaving hundreds or thousands of GPUs idle.
Lossless by Design
InfiniBand RDMA uses credit-based flow control. A receiver advertises buffer credits to a sender. The sender transmits only when credits are available. Packets are never dropped due to buffer overrun. This is different from Ethernet, which is lossy by default and requires PFC and ECN to become lossless for RoCEv2.
That lossless behavior keeps tail latency predictable. In AI training, p99 latency matters more than average latency. One slow packet can bottleneck the whole collective.
Sub-Microsecond Latency
InfiniBand RDMA NDR links deliver small-message latencies around 0.6–0.9 microseconds for RDMA WRITE operations. HDR is roughly 0.8–1.1 microseconds. Even well-tuned 400GbE RoCEv2 usually sits in the 2–5 microsecond range. For tightly coupled collectives, that gap compounds.
CPU Offload
Because the HCA handles transport processing, the host CPU isn’t interrupted for every packet. In a large training cluster, that can free dozens of CPU cores per node for data loading, preprocessing, or checkpointing.
SHARP: All-Reduce in the Switch
SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) offloads collective operations into the switch silicon. In a standard all-reduce, every GPU sends its gradients to every other GPU, which takes O(log N) round-trips. With SHARP enabled on Quantum-2 switches, the switch itself sums gradients in-flight. On clusters of 16 or more nodes, that drops the round-trip count toward O(1).
NVIDIA NCCL enables SHARP with NCCL_COLLNET_ENABLE=1. For 100B+ parameter models on 64+ GPUs, this is often the difference between InfiniBand RDMA being worth the premium and not.

GPUDirect RDMA: GPU-to-GPU Without the CPU
InfiniBand RDMA with GPUDirect is where the technology becomes essential for modern AI clusters. Without it, GPU memory has to copy through host DRAM before the NIC can send it.
The path without GPUDirect looks like this:
GPU HBM → PCIe → CPU DRAM → PCIe → NIC → network → NIC → CPU DRAM → PCIe → GPU HBM
That is two extra copies through system memory and significant CPU involvement. For a 100 GB gradient tensor, the CPU copy alone can add seconds per step.
With GPUDirect RDMA, the NIC DMAs directly into GPU High-Bandwidth Memory (HBM):
GPU HBM → PCIe → NIC → network → NIC → PCIe → GPU HBM
The CPU is removed from the data path entirely.
Requirements
GPUDirect RDMA needs several things to work well:
A supported GPU and HCA. ConnectX-7 and newer HCAs support it.
The nvidia-peermem kernel module (replaced the older nv_peer_mem).
A compatible NVIDIA OFED or rdma-core stack.
GPU and NIC under the same PCIe Root Complex for best performance. Crossing CPU sockets or root complexes adds latency and reduces bandwidth.
The PCIe topology is the detail most teams overlook. A server with GPUs on one CPU socket and NICs on another will run GPUDirect RDMA, but not at full speed. Always verify topology with nvidia-smi topo -m before finalizing rack layout.
InfiniBand RDMA vs RoCEv2
RoCEv2 runs RDMA semantics over standard UDP/IP Ethernet. The same ConnectX-7 HCA can operate in InfiniBand mode or Ethernet/RoCE mode. The choice is largely made at driver configuration time.
| Factor | InfiniBand RDMA | RoCEv2 |
| Latency (p50, small messages) | ~0.6–0.9 µs | ~2–5 µs, down to ~1.5 µs on 800GbE |
| Lossless mechanism | Native credit-based | PFC + ECN (must be configured) |
| Switch vendors | NVIDIA/Mellanox primarily | Arista, Cisco, Juniper, Broadcom |
| Operator expertise | Specialized IB admins | Standard Ethernet skills |
| Multi-tenancy | Limited | EVPN-VXLAN overlays |
| Cost | Higher | 20–30% lower typically |
| SHARP support | Yes | No |
RoCEv2 can deliver 85–95% of InfiniBand training throughput when tuned properly. Meta’s 24,000-GPU RoCEv2 cluster for LLaMA 3.1 proved that. But that “when tuned properly” qualifier is doing a lot of work. Production RoCEv2 requires careful PFC, ECN, DCQCN, and buffer tuning across every switch.
InfiniBand makes sense when every microsecond counts, when SHARP offloads matter, or when the operational team already knows the stack. RoCEv2 makes sense when cost, multi-tenancy, or Ethernet operational familiarity dominate.
For a deeper protocol-level comparison, see our RoCEv2 guide. If you are planning a Quantum-2 deployment, our 800G NDR InfiniBand deployment guide covers the physical layer in detail.
Verifying InfiniBand RDMA Performance
You can’t manage what you don’t measure. After cabling and driver installation, run a short validation workflow to confirm InfiniBand RDMA is performing as expected.
Check port state and speed:
ibstat
ibv_devinfo
You want State: Active, Rate: 400 for NDR, and no excessive error counters.
Measure latency and bandwidth with the perftest tools:
On the server, run:
ib_write_lat -d mlx5_0
On the client, run:
ib_write_lat -d mlx5_0 <server_ip>
For bandwidth, run the same two-machine setup:
On the server, run:
ib_write_bw -d mlx5_0 –report_gbits
On the client, run:
ib_write_bw -d mlx5_0 <server_ip> –report_gbits
Typical small-message WRITE latencies by generation:
| Generation | Link Speed | Typical WRITE Latency |
| EDR | 100 Gb/s | ~1.0–1.3 µs |
| HDR | 200 Gb/s | ~0.8–1.1 µs |
| NDR | 400 Gb/s | ~0.6–0.9 µs |
Note that perftest reports half round-trip for WRITE latency. Bandwidth should approach wire rate for large messages; if it does not, check PCIe topology, CPU frequency, and interrupt affinity.
Common RDMA Pitfalls
Even with good hardware, InfiniBand RDMA deployments go sideways. Here are the issues we see most often.

Wrong cables and connectors. NDR optical links need MPO-12 APC connectors, the green ones, not UPC. Polarity must follow Method B. A dirty connector or reversed polarity can raise the bit error rate enough to silently degrade RDMA performance while the link still shows ACTIVE.
Memory registration bottlenecks. Registering and deregistering buffers for every transfer adds latency. Most high-performance applications register a pool of buffers once and reuse them.
PCIe topology mismatches. GPUDirect RDMA performance drops sharply when the GPU and NIC are not under the same root complex. Check nvidia-smi topo -m before locking the rack design.
QP state errors. A QP must transition RESET → INIT → RTR → RTS in order. Skip a step or pass wrong parameters and the QP lands in ERR state. Recovery means moving it back to RESET and starting over.
Firmware and driver drift. Mismatched OFED versions or outdated HCA firmware cause intermittent RDMA drops. Standardize on one validated stack across the cluster.
The physical layer is usually the culprit. A field engineer once told us: networking problems don’t announce themselves with a clean outage. They hide inside a link that looks fine but performs badly.
FAQ
Do I need InfiniBand for distributed AI training?
Not always. For 8–32 GPUs training models under 70B parameters, well-tuned RoCEv2 or Spectrum-X is usually sufficient. For 64+ GPUs training 100B+ parameter models, InfiniBand RDMA’s lower latency and SHARP offload typically pay off.
What cable do I need for NDR InfiniBand RDMA?
Short distances up to 3 meters can use DAC cables. Mid-range uses AOC. Longer runs need optical transceivers with MPO-12 APC connectors and Method B polarity. Our InfiniBand cable guide breaks down the options.
Is memory registration always required?
For traditional InfiniBand RDMA, yes, buffers must be pinned and keyed. On-Demand Paging can automate registration, but it adds complexity and can cause latency spikes. Most performance-critical code pins buffer pools.
What is the difference between RDMA READ and RDMA WRITE?
RDMA READ pulls data from remote memory to local memory. RDMA WRITE pushes data from local memory to remote memory. Both are one-sided: the remote CPU is not involved.
How do I know GPUDirect RDMA is working?
To confirm InfiniBand RDMA with GPUDirect is working, run ib_write_bw or ib_write_lat with the –use_cuda=<gpu_id> flag. Bandwidth should approach the GPU-NIC PCIe bandwidth, and the CPU should not spike during the transfer. Also verify with nvidia-smi topo -m that GPU and NIC share a root complex.
Conclusion
InfiniBand RDMA removes the CPU and kernel from the data path through kernel bypass, zero-copy transfers, and NIC offload. For AI training clusters, that translates to predictable sub-microsecond latency, lossless behavior, and the ability to offload collectives through SHARP. GPUDirect RDMA extends those gains by letting the NIC DMA directly into GPU memory.
The technology is not plug-and-play. Verbs programming, memory registration, QP state management, and physical-layer details all demand attention. But for workloads where network stalls waste GPU time at scale, InfiniBand RDMA remains the most deterministic option available.
If you are building or expanding an InfiniBand fabric, the cabling and optics are where many deployments win or lose. FiberMall supplies 800G NDR InfiniBand modules, 1.6T OSFP InfiniBand transceivers, and InfiniBand-compatible cables tested for Quantum-2, ConnectX-7, and HGX platforms. Contact our engineering team for a quote or compatibility check.
Related Products:
-
NVIDIA MMA4Z00-NS400 Compatible 400G OSFP SR4 Flat Top PAM4 850nm 30m on OM3/50m on OM4 MTP/MPO-12 Multimode FEC Optical Transceiver Module
$400.00
-
NVIDIA MMS4X00-NS400 Compatible 400G OSFP DR4 Flat Top PAM4 1310nm MTP/MPO-12 500m SMF FEC Optical Transceiver Module
$450.00
-
NVIDIA MMA1Z00-NS400 Compatible 400G QSFP112 VR4 PAM4 850nm 50m MTP/MPO-12 OM4 FEC Optical Transceiver Module
$385.00
-
NVIDIA MMS1X00-NS400 Compatible 400G NDR QSFP112 DR4 PAM4 1310nm 500m MPO-12 with FEC Optical Transceiver Module
$500.00
-
NVIDIA MCP7Y70-H001 Compatible 1m (3ft) 400G Twin-port 2x200G OSFP to 4x100G QSFP56 Passive Breakout Direct Attach Copper Cable
$120.00
-
NVIDIA MCP7Y60-H001 Compatible 1m (3ft) 400G OSFP to 2x200G QSFP56 Passive Direct Attach Cable
$99.00
-
NVIDIA MFA7U10-H003 Compatible 3m (10ft) 400G OSFP to 2x200G QSFP56 twin port HDR Breakout Active Optical Cable
$750.00
-
NVIDIA MMA4Z00-NS Compatible 800GBASE 2 x SR4/SR8 OSFP PAM4 850nm 100m DOM Dual MPO-12 MMF Optical Transceiver Module
$550.00
-
NVIDIA MMA4Z00-NS-FLT Compatible 800GBASE 2 x SR4/SR8 OSFP RHS/Flat Top PAM4 850nm 100m DOM Dual MPO-12 MMF Optical Transceiver Module
$600.00
-
NVIDIA MMS4X00-NM Compatible 800GBASE 2 x DR4/DR8 OSFP IHS/Closed Finned Top PAM4 1310nm 500m DOM Dual MTP/MPO-12 SMF Optical Transceiver Module
$600.00
-
NVIDIA MMS4X00-NM-FLT Compatible 800GBASE 2 x DR4/DR8 OSFP Flat Top PAM4 1310nm 500m DOM Dual MTP/MPO-12 SMF Optical Transceiver Module
$650.00
-
NVIDIA MMS4X50-NM Compatible 800G 2x FR4 OSFP IHS/Closed Finned Top PAM4 1310nm 2km DOM Dual Duplex LC SMF InfiniBand NDR Optical Transceiver Module
$1000.00
-
NVIDIA MCP7Y00-N001 Compatible 1m (3ft) 800Gb Twin-port OSFP to 2x400G OSFP InfiniBand NDR Breakout Direct Attach Copper Cable
$160.00
-
NVIDIA MCA7J60-N004 Compatible 4m (13ft) 800G Twin-port OSFP to 2x400G OSFP InfiniBand NDR Breakout Active Copper Cable
$800.00
-
NVIDIA MMS4A00 (980-9IAH1-00XM00) Compatible 1.6T 2 x DR4/DR8 OSFP224 PAM4 1311nm 500m IHS/Finned Top Dual MPO-12 SMF Optical Transceiver Module
$1500.00
-
NVIDIA MMS4A50 Compatible 1.6T 2xFR4/FR8 OSFP224 PAM4 1310nm 2km IHS/Finned Top Dual Duplex LC SMF Optical Transceiver Module
$1800.00
-
NVIDIA MMS4A00-RHS Compatible 1.6T 2xDR4/DR8 OSFP224 PAM4 1311nm 500m RHS/Flat Top Dual MPO-12/APC InfiniBand XDR SMF Optical Transceiver Module
$2000.00
-
NVIDIA MCA7K20-X001 Compatible 1m (3ft) Twin-port 2x800Gb/s OSFP224 IHS/Finned Top to 4x400Gb/s OSFP224 RHS/Flat Top InfiniBand XDR Active Copper Splitter Cable
$2059.00
Related Posts
- InfiniBand vs Ethernet: Which Network Should Your AI Cluster Use?
- InfiniBand Generations: SDR to GDR Speed Chart (2026)
- InfiniBand vs RoCEv2: AI Data Center Networking Guide (2026)
- InfiniBand Troubleshooting: A Step-by-Step Guide for 2026
- What Is InfiniBand? The Complete Guide to High-Performance AI Networking
