{"id":19686,"date":"2026-06-29T08:03:47","date_gmt":"2026-06-29T08:03:47","guid":{"rendered":"https:\/\/www.fibermall.com\/blog\/?p=19686"},"modified":"2026-07-13T09:47:33","modified_gmt":"2026-07-13T09:47:33","slug":"infiniband-rdma-explained-how-gpu-clusters-communicate-at-microsecond-scale","status":"publish","type":"post","link":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm","title":{"rendered":"InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale"},"content":{"rendered":"\n<p>A single 1% packet-drop rate on a 1,024-GPU H100 cluster can burn roughly $250,000 per week in idle compute. That isn&#8217;t a theory. It&#8217;s why hyperscalers and HPC labs obsess over lossless, low-latency interconnects. InfiniBand RDMA is the technology that makes those interconnects possible.<\/p>\n\n\n\n<p>Remote Direct Memory Access (RDMA) lets one server read from or write to another server&#8217;s memory without waking either CPU or copying data through the operating system kernel. InfiniBand was built specifically for this. While Ethernet with RoCEv2 can run RDMA over a familiar stack, InfiniBand RDMA offers native losslessness and sub-microsecond latency that remains hard to match.<\/p>\n\n\n\n<p>In this article, we&#8217;ll walk through how InfiniBand RDMA actually works: verbs, queue pairs, memory regions, and the connection setup flow. Then we&#8217;ll connect it to AI training, GPUDirect RDMA, and the physical layer that often gets underestimated.<\/p>\n\n\n\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_76 ez-toc-wrap-left counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<span class=\"ez-toc-title-toggle\"><a href=\"#\" class=\"ez-toc-pull-right ez-toc-btn ez-toc-btn-xs ez-toc-btn-default ez-toc-toggle\" aria-label=\"Toggle Table of Content\"><span class=\"ez-toc-js-icon-con\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/span><\/a><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#What_Is_InfiniBand_RDMA\" >What Is InfiniBand RDMA?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#How_InfiniBand_RDMA_Works_The_Software_Stack\" >How InfiniBand RDMA Works: The Software Stack<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Verbs_The_RDMA_API\" >Verbs: The RDMA API<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Queue_Pairs_The_Communication_Endpoint\" >Queue Pairs: The Communication Endpoint<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Memory_Regions_Pinned_and_Keyed\" >Memory Regions: Pinned and Keyed<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Protection_Domains_and_Completion_Queues\" >Protection Domains and Completion Queues<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Connection_Setup_Flow\" >Connection Setup Flow<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#RDMA_Operations_SENDRECEIVE_READ_WRITE\" >RDMA Operations: SEND\/RECEIVE, READ, WRITE<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Why_InfiniBand_RDMA_Dominates_AI_Training\" >Why InfiniBand RDMA Dominates AI Training<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Lossless_by_Design\" >Lossless by Design<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Sub-Microsecond_Latency\" >Sub-Microsecond Latency<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#CPU_Offload\" >CPU Offload<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-13\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#SHARP_All-Reduce_in_the_Switch\" >SHARP: All-Reduce in the Switch<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-14\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#GPUDirect_RDMA_GPU-to-GPU_Without_the_CPU\" >GPUDirect RDMA: GPU-to-GPU Without the CPU<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-15\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Requirements\" >Requirements<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-16\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#InfiniBand_RDMA_vs_RoCEv2\" >InfiniBand RDMA vs RoCEv2<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-17\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Verifying_InfiniBand_RDMA_Performance\" >Verifying InfiniBand RDMA Performance<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-18\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Common_RDMA_Pitfalls\" >Common RDMA Pitfalls<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-19\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#FAQ\" >FAQ<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-20\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Do_I_need_InfiniBand_for_distributed_AI_training\" >Do I need InfiniBand for distributed AI training?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-21\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#What_cable_do_I_need_for_NDR_InfiniBand_RDMA\" >What cable do I need for NDR InfiniBand RDMA?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-22\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Is_memory_registration_always_required\" >Is memory registration always required?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-23\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#What_is_the_difference_between_RDMA_READ_and_RDMA_WRITE\" >What is the difference between RDMA READ and RDMA WRITE?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-24\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#How_do_I_know_GPUDirect_RDMA_is_working\" >How do I know GPUDirect RDMA is working?<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-25\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\/#Conclusion\" >Conclusion<\/a><\/li><\/ul><\/nav><\/div>\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"What_Is_InfiniBand_RDMA\"><\/span><strong>What Is InfiniBand RDMA?<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>InfiniBand RDMA is a networking technology that moves data directly from one application&#8217;s memory to another application&#8217;s memory across a switched fabric. The CPU doesn&#8217;t copy buffers. The kernel doesn&#8217;t process packets. The Host Channel Adapter (HCA), such as NVIDIA ConnectX-7, handles segmentation, reliability, flow control, and ordering in silicon.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img fetchpriority=\"high\" decoding=\"async\" width=\"800\" height=\"436\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/What-Is-InfiniBand-RDMA.png\" alt=\"What Is InfiniBand RDMA\" class=\"wp-image-19687\" style=\"width:800px\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/What-Is-InfiniBand-RDMA.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/What-Is-InfiniBand-RDMA-300x164.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/What-Is-InfiniBand-RDMA-768x419.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>Traditional TCP\/IP follows a longer path. Data copies from user space to kernel socket buffers, then to the NIC, then across the wire. The receiver reverses every step. Each copy and context switch adds latency and consumes CPU cycles. RDMA removes most of that overhead.<\/p>\n\n\n\n<p>The result is three core advantages:<\/p>\n\n\n\n<p><strong>Kernel bypass<\/strong>: Applications post work requests directly to the NIC through user-space libraries.<\/p>\n\n\n\n<p><strong>Zero-copy<\/strong>: Data stays in pinned user buffers; it never transits kernel socket buffers.<\/p>\n\n\n\n<p><strong>NIC offload<\/strong>: Transport logic runs on the HCA, not the host CPU.<\/p>\n\n\n\n<p>For AI clusters, this matters because training workloads generate enormous east-west traffic. A single all-reduce step across a few hundred GPUs can move terabytes of gradients. InfiniBand RDMA keeps that traffic off the CPU and out of the kernel.<\/p>\n\n\n\n<p>These three properties define InfiniBand RDMA and separate it from conventional networking.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"How_InfiniBand_RDMA_Works_The_Software_Stack\"><\/span><strong>How InfiniBand RDMA Works: The Software Stack<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>Most competitors stop at &#8220;RDMA bypasses the kernel.&#8221; That is true, but it skips the part that actually determines whether your code works. InfiniBand RDMA has a distinct software model built around verbs, queue pairs, and memory regions.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Verbs_The_RDMA_API\"><\/span><strong>Verbs: The RDMA API<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>Verbs are the operations an application uses to talk to the HCA. They are defined by the InfiniBand Architecture specification and implemented in libraries like libibverbs&nbsp;and rdma-core. A verb describes an action: post a send request, create a queue pair, register memory, modify a queue-pair state.<\/p>\n\n\n\n<p>The verbs API is lower level than sockets. There is no implicit buffering and no standard blocking send(). The application manages memory, queues, and completions explicitly. This is the interface user-space code uses to drive InfiniBand RDMA.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Queue_Pairs_The_Communication_Endpoint\"><\/span><strong>Queue Pairs: The Communication Endpoint<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>A Queue Pair (QP) is the core communication unit in InfiniBand RDMA. Every QP contains a Send Queue (SQ) and a Receive Queue (RQ). The application posts Work Requests (WRs) to these queues. The HCA consumes them and, when finished, writes Work Completions (WCs) to a Completion Queue (CQ).<\/p>\n\n\n\n<p>QPs support different transport services. Reliable Connection (RC) is the most common: it guarantees in-order delivery and uses acknowledgments. Unreliable Datagram (UD) is lighter and useful for some MPI patterns. Shared Receive Queues let multiple QPs share one RQ to save memory.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" width=\"800\" height=\"436\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/Queue-Pairs-The-Communication-Endpoint.png\" alt=\"Queue Pairs The Communication Endpoint\" class=\"wp-image-19688\" style=\"width:800px\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/Queue-Pairs-The-Communication-Endpoint.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/Queue-Pairs-The-Communication-Endpoint-300x164.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/Queue-Pairs-The-Communication-Endpoint-768x419.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Memory_Regions_Pinned_and_Keyed\"><\/span><strong>Memory Regions: Pinned and Keyed<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>Before the HCA can touch application memory, that memory must be registered as a Memory Region (MR). Registration pins the pages so they cannot be swapped out and builds a virtual-to-physical mapping the NIC can use.<\/p>\n\n\n\n<p>Each MR gets two keys: an L_Key for local access and an R_Key for remote access. For RDMA READ or WRITE, the initiator must know the remote virtual address and the remote R_Key. That information is typically exchanged over a TCP socket or similar out-of-band channel before RDMA traffic starts.<\/p>\n\n\n\n<p>Memory registration has overhead. Pinning large buffers takes time and consumes kernel resources. Some newer systems use On-Demand Paging (ODP) to avoid explicit registration, but ODP can introduce latency spikes if page mappings are not ready.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Protection_Domains_and_Completion_Queues\"><\/span><strong>Protection Domains and Completion Queues<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>A Protection Domain (PD) groups QPs and MRs that belong to the same address-space context. QPs and MRs must share a PD to interact. Completion Queues collect notifications from multiple QPs, and the application can poll them or wait for events.<\/p>\n\n\n\n<p>Polling is faster but burns CPU. Event-driven notification saves CPU but adds latency. Most high-performance workloads poll.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Connection_Setup_Flow\"><\/span><strong>Connection Setup Flow<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>A QP starts in RESET. The application moves it through INIT, then RTR (Ready to Receive), then RTS (Ready to Send). Only in RTS can both sides initiate sends. Each transition requires correct parameters: Local Identifier (LID), Queue Pair Number (QPN), Packet Sequence Number (PSN), and path MTU. Get one wrong and the QP throws a syndrome error that can take time to trace.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"RDMA_Operations_SENDRECEIVE_READ_WRITE\"><\/span><strong>RDMA Operations: SEND\/RECEIVE, READ, WRITE<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>InfiniBand RDMA supports several transfer semantics. Choosing the right one affects both performance and complexity.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Operation<\/strong><\/td><td><strong>Sidedness<\/strong><\/td><td><strong>Remote CPU Involved?<\/strong><\/td><td><strong>Best For<\/strong><\/td><\/tr><tr><td>SEND \/ RECEIVE<\/td><td>Two-sided<\/td><td>Yes, receiver posts buffers<\/td><td>MPI messages, dynamic communication<\/td><\/tr><tr><td>RDMA READ<\/td><td>One-sided<\/td><td>No, initiator pulls data<\/td><td>Metadata lookups, polling reads<\/td><\/tr><tr><td>RDMA WRITE<\/td><td>One-sided<\/td><td>No, initiator pushes data<\/td><td>Bulk transfers, checkpointing<\/td><\/tr><tr><td>Atomic<\/td><td>One-sided<\/td><td>No, hardware does RMW<\/td><td>Locks, counters<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>SEND\/RECEIVE is conceptually closest to messaging. The receiver must pre-post receive buffers. If a SEND arrives with no posted receive, the QP typically enters an error state.<\/p>\n\n\n\n<p>RDMA READ and WRITE are true one-sided operations. The remote CPU does not participate. For a WRITE, the initiator supplies a remote address and R_Key; the remote HCA DMAs data directly into that memory. This is where GPUDirect RDMA becomes possible.<\/p>\n\n\n\n<p>Picking the right operation is a central part of InfiniBand RDMA tuning.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Why_InfiniBand_RDMA_Dominates_AI_Training\"><\/span><strong>Why InfiniBand RDMA Dominates AI Training<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>AI training is unusually sensitive to network behavior. Distributed SGD requires frequent gradient synchronization. A delayed packet at one GPU can stall the entire all-reduce, leaving hundreds or thousands of GPUs idle.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Lossless_by_Design\"><\/span><strong>Lossless by Design<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>InfiniBand RDMA uses credit-based flow control. A receiver advertises buffer credits to a sender. The sender transmits only when credits are available. Packets are never dropped due to buffer overrun. This is different from Ethernet, which is lossy by default and requires PFC and ECN to become lossless for RoCEv2.<\/p>\n\n\n\n<p>That lossless behavior keeps tail latency predictable. In AI training, p99 latency matters more than average latency. One slow packet can bottleneck the whole collective.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Sub-Microsecond_Latency\"><\/span><strong>Sub-Microsecond Latency<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>InfiniBand RDMA NDR links deliver small-message latencies around 0.6\u20130.9 microseconds for RDMA WRITE operations. HDR is roughly 0.8\u20131.1 microseconds. Even well-tuned 400GbE RoCEv2 usually sits in the 2\u20135 microsecond range. For tightly coupled collectives, that gap compounds.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"CPU_Offload\"><\/span><strong>CPU Offload<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>Because the HCA handles transport processing, the host CPU isn&#8217;t interrupted for every packet. In a large training cluster, that can free dozens of CPU cores per node for data loading, preprocessing, or checkpointing.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"SHARP_All-Reduce_in_the_Switch\"><\/span><strong>SHARP: All-Reduce in the Switch<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) offloads collective operations into the switch silicon. In a standard all-reduce, every GPU sends its gradients to every other GPU, which takes O(log N) round-trips. With SHARP enabled on Quantum-2 switches, the switch itself sums gradients in-flight. On clusters of 16 or more nodes, that drops the round-trip count toward O(1).<\/p>\n\n\n\n<p>NVIDIA NCCL enables SHARP with NCCL_COLLNET_ENABLE=1. For 100B+ parameter models on 64+ GPUs, this is often the difference between InfiniBand RDMA being worth the premium and not.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" width=\"800\" height=\"436\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/GPUDirect-RDMA-GPU-to-GPU-Without-the-CPU.png\" alt=\"GPUDirect RDMA GPU-to-GPU Without the CPU\" class=\"wp-image-19689\" style=\"width:800px\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/GPUDirect-RDMA-GPU-to-GPU-Without-the-CPU.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/GPUDirect-RDMA-GPU-to-GPU-Without-the-CPU-300x164.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/GPUDirect-RDMA-GPU-to-GPU-Without-the-CPU-768x419.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"GPUDirect_RDMA_GPU-to-GPU_Without_the_CPU\"><\/span><strong>GPUDirect RDMA: GPU-to-GPU Without the CPU<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>InfiniBand RDMA with GPUDirect is where the technology becomes essential for modern AI clusters. Without it, GPU memory has to copy through host DRAM before the NIC can send it.<\/p>\n\n\n\n<p>The path without GPUDirect looks like this:<\/p>\n\n\n\n<p>GPU HBM \u2192 PCIe \u2192 CPU DRAM \u2192 PCIe \u2192 NIC \u2192 network \u2192 NIC \u2192 CPU DRAM \u2192 PCIe \u2192 GPU HBM<\/p>\n\n\n\n<p>That is two extra copies through system memory and significant CPU involvement. For a 100 GB gradient tensor, the CPU copy alone can add seconds per step.<\/p>\n\n\n\n<p>With GPUDirect RDMA, the NIC DMAs directly into GPU High-Bandwidth Memory (HBM):<\/p>\n\n\n\n<p>GPU HBM \u2192 PCIe \u2192 NIC \u2192 network \u2192 NIC \u2192 PCIe \u2192 GPU HBM<\/p>\n\n\n\n<p>The CPU is removed from the data path entirely.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Requirements\"><\/span><strong>Requirements<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>GPUDirect RDMA needs several things to work well:<\/p>\n\n\n\n<p>A supported GPU and HCA. ConnectX-7 and newer HCAs support it.<\/p>\n\n\n\n<p>The nvidia-peermem&nbsp;kernel module (replaced the older nv_peer_mem).<\/p>\n\n\n\n<p>A compatible NVIDIA OFED or rdma-core&nbsp;stack.<\/p>\n\n\n\n<p>GPU and NIC under the same PCIe Root Complex for best performance. Crossing CPU sockets or root complexes adds latency and reduces bandwidth.<\/p>\n\n\n\n<p>The PCIe topology is the detail most teams overlook. A server with GPUs on one CPU socket and NICs on another will run GPUDirect RDMA, but not at full speed. Always verify topology with nvidia-smi topo -m&nbsp;before finalizing rack layout.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"InfiniBand_RDMA_vs_RoCEv2\"><\/span><strong>InfiniBand RDMA vs RoCEv2<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>RoCEv2 runs RDMA semantics over standard UDP\/IP Ethernet. The same ConnectX-7 HCA can operate in InfiniBand mode or Ethernet\/RoCE mode. The choice is largely made at driver configuration time.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Factor<\/strong><\/td><td><strong>InfiniBand RDMA<\/strong><\/td><td><strong>RoCEv2<\/strong><\/td><\/tr><tr><td>Latency (p50, small messages)<\/td><td>~0.6\u20130.9 \u00b5s<\/td><td>~2\u20135 \u00b5s, down to ~1.5 \u00b5s on 800GbE<\/td><\/tr><tr><td>Lossless mechanism<\/td><td>Native credit-based<\/td><td>PFC + ECN (must be configured)<\/td><\/tr><tr><td>Switch vendors<\/td><td>NVIDIA\/Mellanox primarily<\/td><td>Arista, Cisco, Juniper, Broadcom<\/td><\/tr><tr><td>Operator expertise<\/td><td>Specialized IB admins<\/td><td>Standard Ethernet skills<\/td><\/tr><tr><td>Multi-tenancy<\/td><td>Limited<\/td><td>EVPN-VXLAN overlays<\/td><\/tr><tr><td>Cost<\/td><td>Higher<\/td><td>20\u201330% lower typically<\/td><\/tr><tr><td>SHARP support<\/td><td>Yes<\/td><td>No<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>RoCEv2 can deliver 85\u201395% of InfiniBand training throughput when tuned properly. Meta&#8217;s 24,000-GPU RoCEv2 cluster for LLaMA 3.1 proved that. But that &#8220;when tuned properly&#8221; qualifier is doing a lot of work. Production RoCEv2 requires careful PFC, ECN, DCQCN, and buffer tuning across every switch.<\/p>\n\n\n\n<p>InfiniBand makes sense when every microsecond counts, when SHARP offloads matter, or when the operational team already knows the stack. RoCEv2 makes sense when cost, multi-tenancy, or Ethernet operational familiarity dominate.<\/p>\n\n\n\n<p>For a deeper protocol-level comparison, see our RoCEv2 guide. If you are planning a Quantum-2 deployment, our <a href=\"https:\/\/www.fibermall.com\/blog\/800g-ndr-infiniband-deployment-guide.htm\" target=\"_blank\"><u>800G NDR InfiniBand deployment guide<\/u><\/a>&nbsp;covers the physical layer in detail.<\/p>\n\n\n\n<figure class=\"wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio\"><div class=\"wp-block-embed__wrapper\">\n<div class=\"ast-oembed-container \" style=\"height: 100%;\"><iframe title=\"800G OSFP DR8 InfiniBand Transceiver: Features &amp; Installation Guide | FiberMall\" width=\"500\" height=\"281\" src=\"https:\/\/www.youtube.com\/embed\/p5mCQLPUE8g?feature=oembed\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share\" referrerpolicy=\"strict-origin-when-cross-origin\" allowfullscreen><\/iframe><\/div>\n<\/div><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Verifying_InfiniBand_RDMA_Performance\"><\/span><strong>Verifying InfiniBand RDMA Performance<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>You can&#8217;t manage what you don&#8217;t measure. After cabling and driver installation, run a short validation workflow to confirm InfiniBand RDMA is performing as expected.<\/p>\n\n\n\n<p>Check port state and speed:<\/p>\n\n\n\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;ibstat &nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;ibv_devinfo &nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<\/p>\n\n\n\n<p>You want State: Active, Rate: 400&nbsp;for NDR, and no excessive error counters.<\/p>\n\n\n\n<p>Measure latency and bandwidth with the perftest&nbsp;tools:<\/p>\n\n\n\n<p>On the server, run:<\/p>\n\n\n\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;ib_write_lat -d mlx5_0 &nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<\/p>\n\n\n\n<p>On the client, run:<\/p>\n\n\n\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;ib_write_lat -d mlx5_0 &lt;server_ip&gt; &nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<\/p>\n\n\n\n<p>For bandwidth, run the same two-machine setup:<\/p>\n\n\n\n<p>On the server, run:<\/p>\n\n\n\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;ib_write_bw -d mlx5_0 &#8211;report_gbits &nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<\/p>\n\n\n\n<p>On the client, run:<\/p>\n\n\n\n<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;ib_write_bw -d mlx5_0 &lt;server_ip&gt; &#8211;report_gbits &nbsp;&nbsp;&nbsp;<br>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<\/p>\n\n\n\n<p>Typical small-message WRITE latencies by generation:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Generation<\/strong><\/td><td><strong>Link Speed<\/strong><\/td><td><strong>Typical WRITE Latency<\/strong><\/td><\/tr><tr><td>EDR<\/td><td>100 Gb\/s<\/td><td>~1.0\u20131.3 \u00b5s<\/td><\/tr><tr><td>HDR<\/td><td>200 Gb\/s<\/td><td>~0.8\u20131.1 \u00b5s<\/td><\/tr><tr><td>NDR<\/td><td>400 Gb\/s<\/td><td>~0.6\u20130.9 \u00b5s<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>Note that perftest&nbsp;reports half round-trip for WRITE latency. Bandwidth should approach wire rate for large messages; if it does not, check PCIe topology, CPU frequency, and interrupt affinity.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Common_RDMA_Pitfalls\"><\/span><strong>Common RDMA Pitfalls<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>Even with good hardware, InfiniBand RDMA deployments go sideways. Here are the issues we see most often.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"800\" height=\"483\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/Common-RDMA-Pitfalls.png\" alt=\"Common RDMA Pitfalls\" class=\"wp-image-19690\" style=\"width:800px\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/Common-RDMA-Pitfalls.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/Common-RDMA-Pitfalls-300x181.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/Common-RDMA-Pitfalls-768x464.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p><strong>Wrong cables and connectors.<\/strong>&nbsp;NDR optical links need MPO-12 APC connectors, the green ones, not UPC. Polarity must follow Method B. A dirty connector or reversed polarity can raise the bit error rate enough to silently degrade RDMA performance while the link still shows ACTIVE.<\/p>\n\n\n\n<p><strong>Memory registration bottlenecks.<\/strong>&nbsp;Registering and deregistering buffers for every transfer adds latency. Most high-performance applications register a pool of buffers once and reuse them.<\/p>\n\n\n\n<p><strong>PCIe topology mismatches.<\/strong>&nbsp;GPUDirect RDMA performance drops sharply when the GPU and NIC are not under the same root complex. Check nvidia-smi topo -m&nbsp;before locking the rack design.<\/p>\n\n\n\n<p><strong>QP state errors.<\/strong>&nbsp;A QP must transition RESET \u2192 INIT \u2192 RTR \u2192 RTS in order. Skip a step or pass wrong parameters and the QP lands in ERR state. Recovery means moving it back to RESET and starting over.<\/p>\n\n\n\n<p><strong>Firmware and driver drift.<\/strong>&nbsp;Mismatched OFED versions or outdated HCA firmware cause intermittent RDMA drops. Standardize on one validated stack across the cluster.<\/p>\n\n\n\n<p>The physical layer is usually the culprit. A field engineer once told us: networking problems don&#8217;t announce themselves with a clean outage. They hide inside a link that looks fine but performs badly.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"FAQ\"><\/span><strong>FAQ<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Do_I_need_InfiniBand_for_distributed_AI_training\"><\/span><strong>Do I need InfiniBand for distributed AI training?<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>Not always. For 8\u201332 GPUs training models under 70B parameters, well-tuned RoCEv2 or Spectrum-X is usually sufficient. For 64+ GPUs training 100B+ parameter models, InfiniBand RDMA&#8217;s lower latency and SHARP offload typically pay off.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"What_cable_do_I_need_for_NDR_InfiniBand_RDMA\"><\/span><strong>What cable do I need for NDR InfiniBand RDMA?<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>Short distances up to 3 meters can use DAC cables. Mid-range uses AOC. Longer runs need optical transceivers with MPO-12 APC connectors and Method B polarity. Our <a href=\"https:\/\/www.fibermall.com\/blog\/what-is-infiniband-ai-networking.htm\" target=\"_blank\"><u>InfiniBand cable guide<\/u><\/a>&nbsp;breaks down the options.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Is_memory_registration_always_required\"><\/span><strong>Is memory registration always required?<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>For traditional InfiniBand RDMA, yes, buffers must be pinned and keyed. On-Demand Paging can automate registration, but it adds complexity and can cause latency spikes. Most performance-critical code pins buffer pools.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"What_is_the_difference_between_RDMA_READ_and_RDMA_WRITE\"><\/span><strong>What is the difference between RDMA READ and RDMA WRITE?<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>RDMA READ pulls data from remote memory to local memory. RDMA WRITE pushes data from local memory to remote memory. Both are one-sided: the remote CPU is not involved.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"How_do_I_know_GPUDirect_RDMA_is_working\"><\/span><strong>How do I know GPUDirect RDMA is working?<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h3>\n\n\n\n<p>To confirm InfiniBand RDMA with GPUDirect is working, run ib_write_bw&nbsp;or ib_write_lat&nbsp;with the &#8211;use_cuda=&lt;gpu_id&gt;&nbsp;flag. Bandwidth should approach the GPU-NIC PCIe bandwidth, and the CPU should not spike during the transfer. Also verify with nvidia-smi topo -m&nbsp;that GPU and NIC share a root complex.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Conclusion\"><\/span><strong>Conclusion<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>InfiniBand RDMA removes the CPU and kernel from the data path through kernel bypass, zero-copy transfers, and NIC offload. For AI training clusters, that translates to predictable sub-microsecond latency, lossless behavior, and the ability to offload collectives through SHARP. GPUDirect RDMA extends those gains by letting the NIC DMA directly into GPU memory.<\/p>\n\n\n\n<p>The technology is not plug-and-play. Verbs programming, memory registration, QP state management, and physical-layer details all demand attention. But for workloads where network stalls waste GPU time at scale, InfiniBand RDMA remains the most deterministic option available.<\/p>\n\n\n\n<p>If you are building or expanding an InfiniBand fabric, the cabling and optics are where many deployments win or lose. FiberMall supplies 800G NDR InfiniBand modules, <a href=\"https:\/\/www.fibermall.com\/store-22019-1-6t-osfp.htm\" target=\"_blank\" rel=\"noreferrer noopener\">1.6T OSFP InfiniBand transceivers<\/a>, and InfiniBand-compatible cables tested for Quantum-2, ConnectX-7, and HGX platforms. Contact our engineering team for a quote or compatibility check.<\/p>\n\n\n\n<p><\/p>\n<style>\r\n\r\n        .lwrp.link-whisper-related-posts{\r\n            \r\n            margin-top: 40px;\nmargin-bottom: 30px;\r\n        }\r\n        .lwrp .lwrp-title{\r\n            \r\n            \r\n        }\r\n        .lwrp .lwrp-description{\r\n            \r\n            \r\n\r\n        }\r\n        .lwrp .lwrp-list-container{\r\n        }\r\n        .lwrp .lwrp-list-multi-container{\r\n            display: flex;\r\n        }\r\n        .lwrp .lwrp-list-double{\r\n            width: 48%;\r\n        }\r\n        .lwrp .lwrp-list-triple{\r\n            width: 32%;\r\n        }\r\n        .lwrp .lwrp-list-row-container{\r\n            display: flex;\r\n            justify-content: space-between;\r\n        }\r\n        .lwrp .lwrp-list-row-container .lwrp-list-item{\r\n            width: calc(100% - 20px);\r\n        }\r\n        .lwrp .lwrp-list-item:not(.lwrp-no-posts-message-item){\r\n            \r\n            list-style: decimal;\r\n        }\r\n        .lwrp .lwrp-list-item img{\r\n            max-width: 100%;\r\n            height: auto;\r\n        }\r\n        .lwrp .lwrp-list-item.lwrp-empty-list-item{\r\n            background: initial !important;\r\n        }\r\n        .lwrp .lwrp-list-item .lwrp-list-link .lwrp-list-link-title-text,\r\n        .lwrp .lwrp-list-item .lwrp-list-no-posts-message{\r\n            \r\n                \r\n        }\r\n        @media screen and (max-width: 480px) {\r\n            .lwrp.link-whisper-related-posts{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-title{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-description{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-list-multi-container{\r\n                flex-direction: column;\r\n            }\r\n            .lwrp .lwrp-list-multi-container ul.lwrp-list{\r\n                margin-top: 0px;\r\n                margin-bottom: 0px;\r\n                padding-top: 0px;\r\n                padding-bottom: 0px;\r\n            }\r\n            .lwrp .lwrp-list-double,\r\n            .lwrp .lwrp-list-triple{\r\n                width: 100%;\r\n            }\r\n            .lwrp .lwrp-list-row-container{\r\n                justify-content: initial;\r\n                flex-direction: column;\r\n            }\r\n            .lwrp .lwrp-list-row-container .lwrp-list-item{\r\n                width: 100%;\r\n            }\r\n            .lwrp .lwrp-list-item:not(.lwrp-no-posts-message-item){\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-list-item .lwrp-list-link .lwrp-list-link-title-text,\r\n            .lwrp .lwrp-list-item .lwrp-list-no-posts-message{\r\n                \r\n                    \r\n            }\r\n        }<\/style>\r\n<div id=\"link-whisper-related-posts-widget\" class=\"link-whisper-related-posts lwrp\">\r\n            <h3 class=\"lwrp-title\">Related Posts<\/h3>    \r\n        <div class=\"lwrp-list-container\">\r\n                                            <ul class=\"lwrp-list lwrp-list-single\">\r\n                    <li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/infiniband-vs-ethernet.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">InfiniBand vs Ethernet: Which Network Should Your AI Cluster Use?<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/infiniband-generations.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">InfiniBand Generations: SDR to GDR Speed Chart (2026)<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/infiniband-vs-rocev2.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">InfiniBand vs RoCEv2: AI Data Center Networking Guide (2026)<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/infiniband-troubleshooting-guide.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">InfiniBand Troubleshooting: A Step-by-Step Guide for 2026<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/what-is-infiniband-ai-networking.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">What Is InfiniBand? The Complete Guide to High-Performance AI Networking<\/span><\/a><\/li>                <\/ul>\r\n                        <\/div>\r\n<\/div>","protected":false},"excerpt":{"rendered":"<p>A single 1% packet-drop rate on a 1,024-GPU H100 cluster can burn roughly $250,000 per week in idle compute. That isn&#8217;t a theory. It&#8217;s why hyperscalers and HPC labs obsess over lossless, low-latency interconnects. InfiniBand RDMA is the technology that makes those interconnects possible. Remote Direct Memory Access (RDMA) lets one server read from or [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":19694,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"set","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":"","_wpscppro_dont_share_socialmedia":false,"_wpscppro_custom_social_share_image":0,"_facebook_share_type":"default","_twitter_share_type":"default","_linkedin_share_type":"default","_pinterest_share_type":"default","_linkedin_share_type_page":"default","_instagram_share_type":"default","_medium_share_type":"default","_threads_share_type":"default","_google_business_share_type":"default","_selected_social_profile":[{"id":"skM9ewvR8O","platform":"linkedin","platformKey":0,"name":"Jason Xue","type":"person","thumbnail_url":"https:\/\/media.licdn.com\/dms\/image\/C5603AQErPqKD0j6qBg\/profile-displayphoto-shrink_100_100\/0\/1599138392315?e=1723075200&v=beta&t=joEkh1OeKQ0F-QpAPv4xxQyBdGlHyccIQZauRSs6RvU","share_type":"default"}],"_wpsp_enable_custom_social_template":false,"_wpsp_social_scheduling":{"enabled":false,"datetime":null,"platforms":[],"status":"template_only","dateOption":"today","timeOption":"now","customDays":"","customHours":"","customDate":"","customTime":"","schedulingType":"absolute"},"_wpsp_active_default_template":true},"categories":[2,29],"tags":[],"class_list":["post-19686","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog","category-networking"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v20.13 (Yoast SEO v25.8) - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale - fibermall.com<\/title>\n<meta name=\"description\" content=\"Learn how InfiniBand RDMA enables zero-copy, kernel-bypass networking for AI clusters. Covers verbs, queue pairs, GPUDirect, SHARP, and verification commands.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale\" \/>\n<meta property=\"og:description\" content=\"A single 1% packet-drop rate on a 1,024-GPU H100 cluster can burn roughly $250,000 per week in idle compute. That isn&#039;t a theory. It&#039;s why hyperscalers\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\" \/>\n<meta property=\"og:site_name\" content=\"fibermall.com\" \/>\n<meta property=\"article:published_time\" content=\"2026-06-29T08:03:47+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-13T09:47:33+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/InfiniBand-RDMA.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"900\" \/>\n\t<meta property=\"og:image:height\" content=\"600\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Casey\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Casey\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"12 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#article\",\"isPartOf\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\"},\"author\":{\"name\":\"Casey\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/8a62a1260941aa13aa5900e9ea7e1342\"},\"headline\":\"InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale\",\"datePublished\":\"2026-06-29T08:03:47+00:00\",\"dateModified\":\"2026-07-13T09:47:33+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\"},\"wordCount\":2560,\"publisher\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#primaryimage\"},\"thumbnailUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/InfiniBand-RDMA.jpg\",\"articleSection\":[\"Blog\",\"Networking\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\",\"url\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\",\"name\":\"InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale - fibermall.com\",\"isPartOf\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#primaryimage\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#primaryimage\"},\"thumbnailUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/InfiniBand-RDMA.jpg\",\"datePublished\":\"2026-06-29T08:03:47+00:00\",\"dateModified\":\"2026-07-13T09:47:33+00:00\",\"description\":\"Learn how InfiniBand RDMA enables zero-copy, kernel-bypass networking for AI clusters. Covers verbs, queue pairs, GPUDirect, SHARP, and verification commands.\",\"breadcrumb\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#primaryimage\",\"url\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/InfiniBand-RDMA.jpg\",\"contentUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/InfiniBand-RDMA.jpg\",\"width\":900,\"height\":600,\"caption\":\"InfiniBand RDMA\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/www.fibermall.com\/blog.htm\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#website\",\"url\":\"https:\/\/www.fibermall.com\/blog.htm\/\",\"name\":\"fibermall.com\",\"description\":\"Optical Communication Expert\",\"publisher\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/www.fibermall.com\/blog.htm\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\",\"name\":\"fibermall.com\",\"url\":\"https:\/\/www.fibermall.com\/blog.htm\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg\",\"contentUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg\",\"width\":200,\"height\":67,\"caption\":\"fibermall.com\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/8a62a1260941aa13aa5900e9ea7e1342\",\"name\":\"Casey\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/28096898b03bf5fa03557b33f40633b0fe4fa656c7f28e52669eca8ec4146459?s=96&d=mm&r=g\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/28096898b03bf5fa03557b33f40633b0fe4fa656c7f28e52669eca8ec4146459?s=96&d=mm&r=g\",\"caption\":\"Casey\"},\"description\":\"Expert in access network, PON, GPON, etc.\",\"sameAs\":[\"https:\/\/www.fibermall.com\/\"],\"url\":\"https:\/\/www.fibermall.com\/blog.htm?author=7\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale - fibermall.com","description":"Learn how InfiniBand RDMA enables zero-copy, kernel-bypass networking for AI clusters. Covers verbs, queue pairs, GPUDirect, SHARP, and verification commands.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm","og_locale":"en_US","og_type":"article","og_title":"InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale","og_description":"A single 1% packet-drop rate on a 1,024-GPU H100 cluster can burn roughly $250,000 per week in idle compute. That isn't a theory. It's why hyperscalers","og_url":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm","og_site_name":"fibermall.com","article_published_time":"2026-06-29T08:03:47+00:00","article_modified_time":"2026-07-13T09:47:33+00:00","og_image":[{"width":900,"height":600,"url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/InfiniBand-RDMA.jpg","type":"image\/jpeg"}],"author":"Casey","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Casey","Est. reading time":"12 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#article","isPartOf":{"@id":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm"},"author":{"name":"Casey","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/8a62a1260941aa13aa5900e9ea7e1342"},"headline":"InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale","datePublished":"2026-06-29T08:03:47+00:00","dateModified":"2026-07-13T09:47:33+00:00","mainEntityOfPage":{"@id":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm"},"wordCount":2560,"publisher":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization"},"image":{"@id":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#primaryimage"},"thumbnailUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/InfiniBand-RDMA.jpg","articleSection":["Blog","Networking"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm","url":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm","name":"InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale - fibermall.com","isPartOf":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#primaryimage"},"image":{"@id":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#primaryimage"},"thumbnailUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/InfiniBand-RDMA.jpg","datePublished":"2026-06-29T08:03:47+00:00","dateModified":"2026-07-13T09:47:33+00:00","description":"Learn how InfiniBand RDMA enables zero-copy, kernel-bypass networking for AI clusters. Covers verbs, queue pairs, GPUDirect, SHARP, and verification commands.","breadcrumb":{"@id":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#primaryimage","url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/InfiniBand-RDMA.jpg","contentUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2026\/06\/InfiniBand-RDMA.jpg","width":900,"height":600,"caption":"InfiniBand RDMA"},{"@type":"BreadcrumbList","@id":"https:\/\/www.fibermall.com\/blog\/infiniband-rdma-explained.htm#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.fibermall.com\/blog.htm"},{"@type":"ListItem","position":2,"name":"InfiniBand RDMA Explained: How GPU Clusters Communicate at Microsecond Scale"}]},{"@type":"WebSite","@id":"https:\/\/www.fibermall.com\/blog.htm\/#website","url":"https:\/\/www.fibermall.com\/blog.htm\/","name":"fibermall.com","description":"Optical Communication Expert","publisher":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.fibermall.com\/blog.htm\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization","name":"fibermall.com","url":"https:\/\/www.fibermall.com\/blog.htm\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/","url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg","contentUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg","width":200,"height":67,"caption":"fibermall.com"},"image":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/8a62a1260941aa13aa5900e9ea7e1342","name":"Casey","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/28096898b03bf5fa03557b33f40633b0fe4fa656c7f28e52669eca8ec4146459?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/28096898b03bf5fa03557b33f40633b0fe4fa656c7f28e52669eca8ec4146459?s=96&d=mm&r=g","caption":"Casey"},"description":"Expert in access network, PON, GPON, etc.","sameAs":["https:\/\/www.fibermall.com\/"],"url":"https:\/\/www.fibermall.com\/blog.htm?author=7"}]}},"_links":{"self":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/19686","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=19686"}],"version-history":[{"count":4,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/19686\/revisions"}],"predecessor-version":[{"id":19697,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/19686\/revisions\/19697"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/media\/19694"}],"wp:attachment":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=19686"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=19686"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=19686"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}