{"id":13191,"date":"2024-09-27T09:26:42","date_gmt":"2024-09-27T09:26:42","guid":{"rendered":"https:\/\/www.fibermall.com\/blog\/?p=13191"},"modified":"2024-09-27T09:26:45","modified_gmt":"2024-09-27T09:26:45","slug":"high-performance-gpu-server-hardware-topology-and-cluster-networking","status":"publish","type":"post","link":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm","title":{"rendered":"High-Performance GPU Server Hardware Topology and Cluster Networking"},"content":{"rendered":"\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_76 ez-toc-wrap-left counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<span class=\"ez-toc-title-toggle\"><a href=\"#\" class=\"ez-toc-pull-right ez-toc-btn ez-toc-btn-xs ez-toc-btn-default ez-toc-toggle\" aria-label=\"Toggle Table of Content\"><span class=\"ez-toc-js-icon-con\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/span><\/a><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\/#Terminology_and_Basics\" >Terminology and Basics<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\/#Typical_8A1008A800_Host\" >Typical 8A100\/8A800 Host<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\/#Typical_8H1008H800_Hosts\" >Typical 8H100\/8H800 Hosts<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\/#Typical_4L40S8L40S_Hosts\" >Typical 4*L40S\/8*L40S Hosts<\/a><\/li><\/ul><\/nav><\/div>\n<h2 class=\"wp-block-heading\" id=\"h-terminology-and-basics\"><span class=\"ez-toc-section\" id=\"Terminology_and_Basics\"><\/span><strong>Terminology and Basics<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>Large model training typically utilizes single-machine, 8-GPU hosts to form clusters. The models include 8*{A100, A800, H100, H800}. Below is the hardware topology of a typical 8*A100 GPU host:<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/the-hardware-topology-of-a-typical-8xA100-GPU-host-1024x518.png\" alt=\"the hardware topology of a typical 8xA100 GPU host\" class=\"wp-image-13194\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/the-hardware-topology-of-a-typical-8xA100-GPU-host-1024x518.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/the-hardware-topology-of-a-typical-8xA100-GPU-host-300x152.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/the-hardware-topology-of-a-typical-8xA100-GPU-host-768x388.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/the-hardware-topology-of-a-typical-8xA100-GPU-host.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p><strong>PCIe Switch Chip<\/strong><strong><\/strong><\/p>\n\n\n\n<p>Devices such as CPUs, memory, storage (NVME), GPUs, and network cards that support PCIe can connect to the PCIe bus or a dedicated PCIe switch chip to achieve interconnectivity.<\/p>\n\n\n\n<p>Currently, there are five generations of PCIe products, with the latest being Gen5.<\/p>\n\n\n\n<p>NVLink<\/p>\n\n\n\n<p>Definition<\/p>\n\n\n\n<p>According to Wikipedia, NVLink is a wire-based serial multi-lane near-range communications link developed by Nvidia. Unlike PCI Express, a device can consist of multiple NVLinks, and devices use mesh networking to communicate instead of a central hub. The protocol was first announced in March 2014 and uses a proprietary high-speed signaling interconnect (NVHS).<\/p>\n\n\n\n<p>In summary, NVLink is a high-speed interconnect method between different GPUs within the same host. It is a short-range communication link that ensures successful packet transmission, offers higher performance, and serves as a replacement for PCIe. It supports multiple lanes, with link bandwidth increasing linearly with the number of lanes. GPUs within the same node are interconnected via NVLink in a full-mesh manner (similar to spine-leaf architecture), utilizing NVIDIA\u2019s proprietary technology.<\/p>\n\n\n\n<p>Evolution: Generations 1\/2\/3\/4<\/p>\n\n\n\n<p>The main differences lie in the number of lanes per NVLink and the bandwidth per lane (the figures provided are bidirectional bandwidths):<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/bidirectional-bandwidths-1024x404.png\" alt=\"bidirectional bandwidths\" class=\"wp-image-13195\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/bidirectional-bandwidths-1024x404.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/bidirectional-bandwidths-300x118.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/bidirectional-bandwidths-768x303.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/bidirectional-bandwidths.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>For example:<\/p>\n\n\n\n<p>A100: 2 lanes\/NVSwitch * 6 NVSwitch * 50GB\/s\/lane = 600GB\/s bidirectional bandwidth (300GB\/s unidirectional). Note: This is the total bandwidth from one GPU to all NVSwitches.<\/p>\n\n\n\n<p>A800: Reduced by 4 lanes, resulting in 8 lanes * 50GB\/s\/lane = 400GB\/s bidirectional bandwidth (200GB\/s unidirectional).<\/p>\n\n\n\n<p>Monitoring<\/p>\n\n\n\n<p>Real-time NVLink bandwidth can be collected based on DCGM metrics.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Real-time-NVLink-bandwidth-can-be-collected-based-on-DCGM-metrics-1024x338.png\" alt=\"Real-time NVLink bandwidth can be collected based on DCGM metrics.\" class=\"wp-image-13196\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Real-time-NVLink-bandwidth-can-be-collected-based-on-DCGM-metrics-1024x338.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Real-time-NVLink-bandwidth-can-be-collected-based-on-DCGM-metrics-300x99.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Real-time-NVLink-bandwidth-can-be-collected-based-on-DCGM-metrics-768x254.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Real-time-NVLink-bandwidth-can-be-collected-based-on-DCGM-metrics.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>NVSwitch<\/p>\n\n\n\n<p>Refer to the diagram below for a typical 8*A100 GPU host hardware topology.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/NVSwitch-is-an-NVIDIA-switch-chip-encapsulated-within-the-GPU-module-1024x518.png\" alt=\"NVSwitch is an NVIDIA switch chip encapsulated within the GPU module\" class=\"wp-image-13197\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/NVSwitch-is-an-NVIDIA-switch-chip-encapsulated-within-the-GPU-module-1024x518.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/NVSwitch-is-an-NVIDIA-switch-chip-encapsulated-within-the-GPU-module-300x152.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/NVSwitch-is-an-NVIDIA-switch-chip-encapsulated-within-the-GPU-module-768x388.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/NVSwitch-is-an-NVIDIA-switch-chip-encapsulated-within-the-GPU-module.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>NVSwitch is an NVIDIA switch chip encapsulated within the GPU module, not an independent external switch.<\/p>\n\n\n\n<p>Below is an image of an actual machine from Inspur. The eight boxes represent the eight A100 GPUs, and the six thick heat sinks on the right cover the NVSwitch chips:<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/an-actual-machine-from-Inspur.png\" alt=\"an actual machine from Inspur.\" class=\"wp-image-13198\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/an-actual-machine-from-Inspur.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/an-actual-machine-from-Inspur-300x200.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/an-actual-machine-from-Inspur-768x513.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>NVLink Switch<\/p>\n\n\n\n<p>Although NVSwitch sounds like a switch, it is actually a switch chip on the GPU module used to connect GPUs within the same host.<\/p>\n\n\n\n<p>In 2022, NVIDIA released this chip as an actual switch called NVLink Switch, designed to connect GPU devices across hosts. The names can be easily confused.<\/p>\n\n\n\n<p>HBM (High Bandwidth Memory)<\/p>\n\n\n\n<p>Origin<\/p>\n\n\n\n<p>Traditionally, GPU memory and regular memory (DDR) are mounted on the motherboard and connected to the processor (CPU, GPU) via PCIe. This creates a speed bottleneck at PCIe, with Gen4 offering 64GB\/s and Gen5 offering 128GB\/s. To overcome this, some GPU manufacturers (not just NVIDIA) stack multiple DDR chips and package them with the GPU. This way, each GPU can interact with its own memory without routing through the PCIe switch chip, significantly increasing speed. This \u201cHigh Bandwidth Memory\u201d is abbreviated as HBM. The HBM market is currently dominated by South Korean companies like SK Hynix and Samsung.<\/p>\n\n\n\n<p>Evolution: HBM 1\/2\/2e\/3\/3e<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Evolution-HBM.png\" alt=\"Evolution HBM\" class=\"wp-image-13199\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Evolution-HBM.png 766w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Evolution-HBM-300x209.png 300w\" sizes=\"(max-width: 766px) 100vw, 766px\" \/><\/figure>\n\n\n\n<p>According to Wikipedia, the AMD MI300X uses a 192GB HBM3 configuration with a bandwidth of 5.2TB\/s. HBM3e is an enhanced version of HBM3, with speeds ranging from 6.4GT\/s to 8GT\/s.<\/p>\n\n\n\n<p>Bandwidth Units<\/p>\n\n\n\n<p>The performance of large-scale GPU training is directly related to data transfer speeds. This involves various links, such as PCIe bandwidth, memory bandwidth, NVLink bandwidth, HBM bandwidth, and network bandwidth. Network bandwidth is typically expressed in bits per second (b\/s) and usually refers to unidirectional (TX\/RX). Other modules\u2019 bandwidth is generally expressed in bytes per second (B\/s) or transactions per second (T\/s) and usually refers to total bidirectional bandwidth. It is important to distinguish and convert these units when comparing bandwidths.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-typical-8a100-8a800-host\"><span class=\"ez-toc-section\" id=\"Typical_8A1008A800_Host\"><\/span><strong>Typical 8A100\/8A800 Host<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>Host Topology: 2-2-4-6-8-8<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>2 CPUs (and their respective memory, NUMA)<\/li>\n\n\n\n<li>2 storage network cards (for accessing distributed storage, in-band management, etc.)<\/li>\n\n\n\n<li>4 PCIe Gen4 Switch chips<\/li>\n\n\n\n<li>6 NVSwitch chips<\/li>\n\n\n\n<li>8 GPUs<\/li>\n\n\n\n<li>8 GPU-dedicated network cards<\/li>\n<\/ul>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology-1024x518.png\" alt=\"Host Topology\" class=\"wp-image-13200\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology-1024x518.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology-300x152.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology-768x388.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>The following diagram provides a more detailed view:<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/The-following-diagram-provides-a-more-detailed-view-799x1024.png\" alt=\"The following diagram provides a more detailed view\" class=\"wp-image-13201\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/The-following-diagram-provides-a-more-detailed-view-799x1024.png 799w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/The-following-diagram-provides-a-more-detailed-view-234x300.png 234w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/The-following-diagram-provides-a-more-detailed-view-768x984.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/The-following-diagram-provides-a-more-detailed-view.png 1080w\" sizes=\"(max-width: 799px) 100vw, 799px\" \/><\/figure>\n\n\n\n<p>Storage Network Cards<\/p>\n\n\n\n<p>These are directly connected to the CPU via PCIe. Their purposes include:<\/p>\n\n\n\n<p>Reading and writing data from distributed storage, such as reading training data and writing checkpoints.<\/p>\n\n\n\n<p>Normal node management, SSH, monitoring, etc.<\/p>\n\n\n\n<p>The official recommendation is to use BF3 DPU, but as long as the bandwidth meets the requirements, any solution will work. For cost-effective networking, use RoCE; for the best performance, use IB.<\/p>\n\n\n\n<p>NVSwitch Fabric: Intra-Node Full-Mesh<\/p>\n\n\n\n<p>The 8 GPUs are connected in a full-mesh configuration via 6 NVSwitch chips, also known as NVSwitch fabric. Each link in the full-mesh has a bandwidth of n * bw-per-nvlink-lane:<\/p>\n\n\n\n<p>For A100 using NVLink3, it is 50GB\/s per lane, so each link in the full-mesh is 12*50GB\/s = 600GB\/s (bidirectional), with 300GB\/s unidirectional.<\/p>\n\n\n\n<p>For A800, which is a reduced version, 12 lanes are reduced to 8 lanes, so each link is 8*50GB\/s = 400GB\/s (bidirectional), with 200GB\/s unidirectional.<\/p>\n\n\n\n<p>Using nvidia-smi topo to View Topology<\/p>\n\n\n\n<p>Below is the actual topology displayed by nvidia-smi on an 8*A800 machine (network cards are bonded in pairs, NIC 0~3 are bonded):<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/the-actual-topology-displayed-by-nvidia-smi-1024x506.png\" alt=\"the actual topology displayed by nvidia-smi\" class=\"wp-image-13202\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/the-actual-topology-displayed-by-nvidia-smi-1024x506.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/the-actual-topology-displayed-by-nvidia-smi-300x148.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/the-actual-topology-displayed-by-nvidia-smi-768x380.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/the-actual-topology-displayed-by-nvidia-smi.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>Between GPUs (top left area): All are NV8, indicating 8 NVLink connections.<\/p>\n\n\n\n<p>Between NICs:<\/p>\n\n\n\n<p>On the same CPU: NODE, indicating no need to cross NUMA but requires crossing PCIe switch chips.<\/p>\n\n\n\n<p>On different CPUs: SYS, indicating the need to cross NUMA.<\/p>\n\n\n\n<p>Between GPUs and NICs:<\/p>\n\n\n\n<p>On the same CPU and under the same PCIe Switch chip: NODE, indicating only crossing PCIe switch chips.<\/p>\n\n\n\n<p>On the same CPU but under different PCIe Switch chips: NODE, indicating crossing PCIe switch chips and PCIe Host Bridge.<\/p>\n\n\n\n<p>On different CPUs: SYS, indicating crossing NUMA, PCIe switch chips, and the longest distance.<\/p>\n\n\n\n<p>GPU Training Cluster Networking: IDC GPU Fabric<\/p>\n\n\n\n<p>GPU Node Interconnection Architecture:<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/GPU-Node-Interconnection-Architecture-1024x430.png\" alt=\"GPU Node Interconnection Architecture\" class=\"wp-image-13203\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/GPU-Node-Interconnection-Architecture-1024x430.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/GPU-Node-Interconnection-Architecture-300x126.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/GPU-Node-Interconnection-Architecture-768x322.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/GPU-Node-Interconnection-Architecture.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>Compute Network:<\/p>\n\n\n\n<p>The GPU network interface cards (NICs) are directly connected to the top-of-rack switches (leaf). These leaf switches are connected in a full-mesh topology to the spine switches, forming an inter-host GPU compute network. The purpose of this network is to facilitate data exchange between GPUs on different nodes. Each GPU is connected to its NIC through a PCIe switch chip: GPU &lt;\u2013&gt; PCIe Switch &lt;\u2013&gt; NIC.<\/p>\n\n\n\n<p>Storage Network:<\/p>\n\n\n\n<p>Two NICs directly connected to the CPU are linked to another network, primarily for data read\/write operations and SSH management.<\/p>\n\n\n\n<p>RoCE vs. InfiniBand:<\/p>\n\n\n\n<p>Both the compute and storage networks require RDMA to achieve the high performance needed for AI. Currently, there are two RDMA options:<\/p>\n\n\n\n<p>RoCEv2: Public cloud providers typically use this network for 8-GPU hosts, such as the CX6 with an 8*100Gbps configuration. It is relatively inexpensive while meeting performance requirements.<\/p>\n\n\n\n<p>InfiniBand (IB): Offers over 20% better performance than RoCEv2 at the same NIC bandwidth but is twice as expensive.<\/p>\n\n\n\n<p>Data Link Bandwidth Bottleneck Analysis:<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Data-Link-Bandwidth-Bottleneck-Analysis-1024x466.png\" alt=\"Data Link Bandwidth Bottleneck Analysis\" class=\"wp-image-13204\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Data-Link-Bandwidth-Bottleneck-Analysis-1024x466.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Data-Link-Bandwidth-Bottleneck-Analysis-300x137.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Data-Link-Bandwidth-Bottleneck-Analysis-768x350.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Data-Link-Bandwidth-Bottleneck-Analysis.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>Key link bandwidths are indicated in the diagram:<\/p>\n\n\n\n<p>Intra-host GPU communication: Utilizes NVLink with a bidirectional bandwidth of 600GB\/s (300GB\/s unidirectional).<\/p>\n\n\n\n<p>Intra-host GPU to NIC communication: Uses PCIe Gen4 switch chips with a bidirectional bandwidth of 64GB\/s (32GB\/s unidirectional).<\/p>\n\n\n\n<p>Inter-host GPU communication: Relies on NICs for data transmission. The mainstream bandwidth for domestic A100\/A800 models is 100Gbps (12.5GB\/s unidirectional), resulting in significantly lower performance compared to intra-host communication.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>200Gbps (25GB\/s): Approaches the unidirectional bandwidth of PCIe Gen4.<\/li>\n\n\n\n<li>400Gbps (50GB\/s): Exceeds the unidirectional bandwidth of PCIe Gen4.<\/li>\n<\/ul>\n\n\n\n<p>Thus, using 400Gbps NICs in these models is not very effective, as 400Gbps requires PCIe Gen5 performance to be fully utilized.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-typical-8h100-8h800-hosts\"><span class=\"ez-toc-section\" id=\"Typical_8H1008H800_Hosts\"><\/span><strong>Typical 8H100\/8H800 Hosts<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>There are two types&nbsp;of GPU Board Form Factor:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>PCIe Gen5<\/li>\n\n\n\n<li>SXM5: Offers higher performance.<\/li>\n<\/ul>\n\n\n\n<p>H100 Chip Layout:<\/p>\n\n\n\n<p>The internal structure of an H100 GPU chip includes:<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/The-internal-structure-of-an-H100-GPU-chip-1024x492.png\" alt=\"The internal structure of an H100 GPU chip\" class=\"wp-image-13205\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/The-internal-structure-of-an-H100-GPU-chip-1024x492.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/The-internal-structure-of-an-H100-GPU-chip-300x144.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/The-internal-structure-of-an-H100-GPU-chip-768x369.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/The-internal-structure-of-an-H100-GPU-chip.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>4nm process technology.<\/p>\n\n\n\n<p>The bottom row consists of 18 Gen4 NVLink lanes with a total bidirectional bandwidth of 900GB\/s (18 lanes * 25GB\/s\/lane).<\/p>\n\n\n\n<p>The middle blue section is the L2 cache.<\/p>\n\n\n\n<p>The sides contain HBM chips, which serve as the GPU memory.<\/p>\n\n\n\n<p>Intra-host Hardware Topology:<\/p>\n\n\n\n<p>Similar to the A100 8-GPU structure, with the following differences:<\/p>\n\n\n\n<p>The number of NVSwitch chips has been reduced from 6 to 4.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Intra-host-Hardware-Topology-1024x659.png\" alt=\"Intra-host Hardware Topology\" class=\"wp-image-13206\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Intra-host-Hardware-Topology-1024x659.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Intra-host-Hardware-Topology-300x193.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Intra-host-Hardware-Topology-768x494.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Intra-host-Hardware-Topology.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>The connection to the CPU has been upgraded from PCIe Gen4 x16 to PCIe Gen5 x16, with a bidirectional bandwidth of 128GB\/s.<\/p>\n\n\n\n<p>Networking:<\/p>\n\n\n\n<p>Similar to the A100, but the standard configuration has been upgraded to 400Gbps CX7 NICs. Otherwise, the bandwidth gap between the PCIe switch and NVLink\/NVSwitch would be even larger.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-typical-4-l40s-8-l40s-hosts\"><span class=\"ez-toc-section\" id=\"Typical_4L40S8L40S_Hosts\"><\/span><strong>Typical 4*L40S\/8*L40S Hosts<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>The L40S is a new generation of cost-effective, multifunctional GPUs set to be released in 2023, positioned as a competitor to the A100. While it is not suitable for training large foundational models (as will be explained later), it is advertised as being capable of handling almost any other task.<\/p>\n\n\n\n<p>Comparison of L40S and A100 Configurations and Features<\/p>\n\n\n\n<p>One of the key features of the L40S is its short time-to-market, meaning the period from order to delivery is much shorter compared to the A100\/A800\/H800. This is due to both technical and non-technical reasons, such as: The removal of FP64 and NVLink.<\/p>\n\n\n\n<p>The use of GDDR6 memory, which does not rely on HBM production capacity (and advanced packaging).<\/p>\n\n\n\n<p>The lower cost is attributed to several factors, which will be detailed later.<\/p>\n\n\n\n<p>The primary cost reduction likely comes from the GPU itself, due to the removal of certain modules and functions or the use of cheaper alternatives.<\/p>\n\n\n\n<p>Savings in the overall system cost, such as the elimination of a layer of PCIe Gen4 switches. Compared to 4x\/8x GPUs, the rest of the system components are almost negligible in cost.<\/p>\n\n\n\n<p>Performance Comparison Between L40S and A100<\/p>\n\n\n\n<p>Below is an official performance comparison:<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Performance-Comparison-Between-L40S-and-A100.png\" alt=\"Performance Comparison Between L40S and A100\" class=\"wp-image-13207\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Performance-Comparison-Between-L40S-and-A100.png 982w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Performance-Comparison-Between-L40S-and-A100-300x232.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Performance-Comparison-Between-L40S-and-A100-768x594.png 768w\" sizes=\"(max-width: 982px) 100vw, 982px\" \/><\/figure>\n\n\n\n<p>Performance: 1.2x to 2x (depending on the specific scenario).<\/p>\n\n\n\n<p>Power consumption: Two L40S units consume roughly the same power as a single A100.<\/p>\n\n\n\n<p>It is important to note that the official recommendation for L40S hosts is a single machine with 4 GPUs rather than 8 (the reasons for this will be explained later). Therefore, comparisons are generally made between two 4L40S units and a single 8A100 unit. Additionally, many performance improvements in various scenarios have a major prerequisite: the network must be a 200Gbps RoCE or IB network, which will be explained next.<\/p>\n\n\n\n<p>L40S System Assembly<\/p>\n\n\n\n<p>Recommended Architecture: 2-2-4<\/p>\n\n\n\n<p>Compared to the A100\u2019s 2-2-4-6-8-8 architecture, the officially recommended L40S GPU host architecture is 2-2-4. The physical topology of a single machine is as follows:<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Recommended-Architecture-2-2-4-1024x430.png\" alt=\"Recommended Architecture 2-2-4\" class=\"wp-image-13208\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Recommended-Architecture-2-2-4-1024x430.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Recommended-Architecture-2-2-4-300x126.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Recommended-Architecture-2-2-4-768x322.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Recommended-Architecture-2-2-4.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>The most noticeable change is the removal of the PCIe switch chip between the CPU and GPU. Both the NIC and GPU are directly connected to the CPU\u2019s built-in PCIe Gen4 x16 (64GB\/s):<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>2 CPUs (NUMA)<\/li>\n\n\n\n<li>2 dual-port CX7 NICs (each NIC 2*200Gbps)<\/li>\n\n\n\n<li>4 L40S GPUs<\/li>\n<\/ul>\n\n\n\n<p>Additionally, only one dual-port storage NIC is provided, directly connected to any one of the CPUs.<\/p>\n\n\n\n<p>This configuration provides each GPU with an average network bandwidth of 200Gbps.<\/p>\n\n\n\n<p>Non-Recommended Architecture: 2-2-8<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Non-Recommended-Architecture-2-2-8.png\" alt=\"Non-Recommended Architecture 2-2-8\" class=\"wp-image-13209\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Non-Recommended-Architecture-2-2-8.png 668w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Non-Recommended-Architecture-2-2-8-300x139.png 300w\" sizes=\"(max-width: 668px) 100vw, 668px\" \/><\/figure>\n\n\n\n<p>As shown, compared to a single machine with 4 GPUs, a single machine with 8 GPUs requires the introduction of two PCIe Gen5 switch chips.<\/p>\n\n\n\n<p>It is said that the current price of a single PCIe Gen5 switch chip is $10,000 (though this is unverified), and a single machine requires 2 chips, making it cost-ineffective.<\/p>\n\n\n\n<p>There is only one manufacturer producing PCIe switches, with limited production capacity and long lead times.<\/p>\n\n\n\n<p>The network bandwidth per GPU is halved.<\/p>\n\n\n\n<p>Networking<\/p>\n\n\n\n<p>The official recommendation is for 4-GPU models, paired with 200Gbps RoCE\/IB networking.<\/p>\n\n\n\n<p><strong>Analysis of Data Link Bandwidth Bottlenecks<\/strong><strong><\/strong><\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Analysis-of-Data-Link-Bandwidth-Bottlenecks-1024x421.png\" alt=\"Analysis of Data Link Bandwidth Bottlenecks\" class=\"wp-image-13210\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Analysis-of-Data-Link-Bandwidth-Bottlenecks-1024x421.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Analysis-of-Data-Link-Bandwidth-Bottlenecks-300x123.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Analysis-of-Data-Link-Bandwidth-Bottlenecks-768x316.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Analysis-of-Data-Link-Bandwidth-Bottlenecks.png 1080w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>Using two L40S GPUs under the same CPU as an example, there are two possible link options:<\/p>\n\n\n\n<ol class=\"wp-block-list\" type=\"1\">\n<li>Direct CPU Processing:<\/li>\n<\/ol>\n\n\n\n<p>Path: GPU0 &lt;\u2013PCIe\u2013&gt; CPU &lt;\u2013PCIe\u2013&gt; GPU1<\/p>\n\n\n\n<p>Bandwidth: PCIe Gen4 x16 with a bidirectional bandwidth of 64GB\/s (32GB\/s unidirectional).<\/p>\n\n\n\n<p>CPU Processing Bottleneck: To be determined.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Bypassing CPU Processing:<\/li>\n<\/ul>\n\n\n\n<p>Path: GPU0 &lt;\u2013PCIe\u2013&gt; NIC &lt;&#8211; RoCE\/IB Switch &#8211;&gt; NIC &lt;\u2013PCIe\u2013&gt; GPU1<\/p>\n\n\n\n<p>Bandwidth: PCIe Gen4 x16 with a bidirectional bandwidth of 64GB\/s (32GB\/s unidirectional).<\/p>\n\n\n\n<p>Average Bandwidth per GPU: Each GPU has a unidirectional 200Gbps network port, equivalent to 25GB\/s.<\/p>\n\n\n\n<p>NCCL Support: The latest version of NCCL is being adapted for the L40S, with the default behavior routing data externally and back.<\/p>\n\n\n\n<p>Although this method appears longer, it is reportedly faster than the first method, provided the NICs and switches are properly configured with a 200Gbps RoCE\/IB network. In this network architecture, with sufficient bandwidth, the communication bandwidth and latency between any two GPUs are consistent, regardless of whether they are within the same machine or under the same CPU. This allows for horizontal scaling of the cluster.<\/p>\n\n\n\n<p>Cost and Performance Considerations:<\/p>\n\n\n\n<p>The cost of GPU machines is reduced. However, for tasks with lower network bandwidth requirements, the cost of NVLINK is effectively transferred to the network. Therefore, a 200Gbps network is essential to fully utilize the performance of multi-GPU training with the L40S.<\/p>\n\n\n\n<p>Bandwidth Bottlenecks in Method Two:<\/p>\n\n\n\n<p>The bandwidth bottleneck between GPUs within the same host is determined by the NIC speed. Even with the recommended 2*CX7 configuration:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>L40S: 200Gbps (unidirectional NIC speed)<\/li>\n\n\n\n<li>A100: 300GB\/s (unidirectional NVLINK3) == 12x200Gbps<\/li>\n\n\n\n<li>A800: 200GB\/s (unidirectional NVLINK3) == 8x200Gbps<\/li>\n<\/ul>\n\n\n\n<p>It is evident that the inter-GPU bandwidth of the L40S is 12 times slower than the A100 NVLINK and 8 times slower than the A800 NVLINK, making it unsuitable for data-intensive foundational model training.<\/p>\n\n\n\n<p>Testing Considerations:<\/p>\n\n\n\n<p>As mentioned, even when testing a single 4-GPU L40S machine, a 200Gbps switch is required to achieve optimal inter-GPU performance.<\/p>\n<style>\r\n\r\n        .lwrp.link-whisper-related-posts{\r\n            \r\n            margin-top: 40px;\nmargin-bottom: 30px;\r\n        }\r\n        .lwrp .lwrp-title{\r\n            \r\n            \r\n        }\r\n        .lwrp .lwrp-description{\r\n            \r\n            \r\n\r\n        }\r\n        .lwrp .lwrp-list-container{\r\n        }\r\n        .lwrp .lwrp-list-multi-container{\r\n            display: flex;\r\n        }\r\n        .lwrp .lwrp-list-double{\r\n            width: 48%;\r\n        }\r\n        .lwrp .lwrp-list-triple{\r\n            width: 32%;\r\n        }\r\n        .lwrp .lwrp-list-row-container{\r\n            display: flex;\r\n            justify-content: space-between;\r\n        }\r\n        .lwrp .lwrp-list-row-container .lwrp-list-item{\r\n            width: calc(100% - 20px);\r\n        }\r\n        .lwrp .lwrp-list-item:not(.lwrp-no-posts-message-item){\r\n            \r\n            list-style: decimal;\r\n        }\r\n        .lwrp .lwrp-list-item img{\r\n            max-width: 100%;\r\n            height: auto;\r\n        }\r\n        .lwrp .lwrp-list-item.lwrp-empty-list-item{\r\n            background: initial !important;\r\n        }\r\n        .lwrp .lwrp-list-item .lwrp-list-link .lwrp-list-link-title-text,\r\n        .lwrp .lwrp-list-item .lwrp-list-no-posts-message{\r\n            \r\n                \r\n        }\r\n        @media screen and (max-width: 480px) {\r\n            .lwrp.link-whisper-related-posts{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-title{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-description{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-list-multi-container{\r\n                flex-direction: column;\r\n            }\r\n            .lwrp .lwrp-list-multi-container ul.lwrp-list{\r\n                margin-top: 0px;\r\n                margin-bottom: 0px;\r\n                padding-top: 0px;\r\n                padding-bottom: 0px;\r\n            }\r\n            .lwrp .lwrp-list-double,\r\n            .lwrp .lwrp-list-triple{\r\n                width: 100%;\r\n            }\r\n            .lwrp .lwrp-list-row-container{\r\n                justify-content: initial;\r\n                flex-direction: column;\r\n            }\r\n            .lwrp .lwrp-list-row-container .lwrp-list-item{\r\n                width: 100%;\r\n            }\r\n            .lwrp .lwrp-list-item:not(.lwrp-no-posts-message-item){\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-list-item .lwrp-list-link .lwrp-list-link-title-text,\r\n            .lwrp .lwrp-list-item .lwrp-list-no-posts-message{\r\n                \r\n                    \r\n            }\r\n        }<\/style>\r\n<div id=\"link-whisper-related-posts-widget\" class=\"link-whisper-related-posts lwrp\">\r\n            <h3 class=\"lwrp-title\">Related Posts<\/h3>    \r\n        <div class=\"lwrp-list-container\">\r\n                                            <ul class=\"lwrp-list lwrp-list-single\">\r\n                    <li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/questions\/when-100-swdm4-srbd-be-used.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">Both the 100G-SWDM4 and the 100G-SRBD Transceivers Support 100G over Duplex Multi-Mode Fiber. When Should Each Transceiver Be Used?<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/24-port-poe-switch.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">Unlock the Potential of Your Network with a 24-Port PoE Switch<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/qsfp-dd-vs-qsfp28.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">Comparing Form Factors: QSFP-DD vs QSFP28 &#8211; Understanding the Differences &#8211; LightOptics\u00ae<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/100g-qsfp28-lr4-psm4-and-cwdm4-modules.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">How to Tell from 100G QSFP28 LR4, PSM4 and CWDM4 Modules Clearly<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/what-is-liquid-cooled-switch.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">What is Liquid Cooled Switch?<\/span><\/a><\/li>                <\/ul>\r\n                        <\/div>\r\n<\/div>","protected":false},"excerpt":{"rendered":"<p>Terminology and Basics Large model training typically utilizes single-machine, 8-GPU hosts to form clusters. The models include 8*{A100, A800, H100, H800}. Below is the hardware topology of a typical 8*A100 GPU host: PCIe Switch Chip Devices such as CPUs, memory, storage (NVME), GPUs, and network cards that support PCIe can connect to the PCIe bus [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":13200,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"set","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":"","_wpscppro_dont_share_socialmedia":false,"_wpscppro_custom_social_share_image":0,"_facebook_share_type":"default","_twitter_share_type":"default","_linkedin_share_type":"default","_pinterest_share_type":"default","_linkedin_share_type_page":"default","_instagram_share_type":"default","_medium_share_type":"","_threads_share_type":"","_google_business_share_type":"","_selected_social_profile":[{"id":"skM9ewvR8O","platform":"linkedin","platformKey":0,"name":"Jason Xue","type":"person","thumbnail_url":"https:\/\/media.licdn.com\/dms\/image\/C5603AQErPqKD0j6qBg\/profile-displayphoto-shrink_100_100\/0\/1599138392315?e=1723075200&v=beta&t=joEkh1OeKQ0F-QpAPv4xxQyBdGlHyccIQZauRSs6RvU","share_type":"default"}],"_wpsp_enable_custom_social_template":false,"_wpsp_social_scheduling":{"enabled":false,"datetime":null,"platforms":[],"status":"template_only","dateOption":"today","timeOption":"now","customDays":"","customHours":"","customDate":"","customTime":"","schedulingType":"absolute"},"_wpsp_active_default_template":true},"categories":[2,29],"tags":[],"class_list":["post-13191","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog","category-networking"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v20.13 (Yoast SEO v25.8) - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>High-Performance GPU Server Hardware Topology and Cluster Networking - fibermall.com<\/title>\n<meta name=\"description\" content=\"NVSwitch is an NVIDIA switch chip encapsulated within the GPU module, not an independent external switch.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"High-Performance GPU Server Hardware Topology and Cluster Networking\" \/>\n<meta property=\"og:description\" content=\"Terminology and Basics Large model training typically utilizes single-machine, 8-GPU hosts to form clusters. The models include 8*{A100, A800, H100,\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\" \/>\n<meta property=\"og:site_name\" content=\"fibermall.com\" \/>\n<meta property=\"article:published_time\" content=\"2024-09-27T09:26:42+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2024-09-27T09:26:45+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1080\" \/>\n\t<meta property=\"og:image:height\" content=\"546\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"FiberMall\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"FiberMall\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"14 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#article\",\"isPartOf\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\"},\"author\":{\"name\":\"FiberMall\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/68e73044a5a439d9e8f42e16adcd0a86\"},\"headline\":\"High-Performance GPU Server Hardware Topology and Cluster Networking\",\"datePublished\":\"2024-09-27T09:26:42+00:00\",\"dateModified\":\"2024-09-27T09:26:45+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\"},\"wordCount\":2255,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#primaryimage\"},\"thumbnailUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png\",\"articleSection\":[\"Blog\",\"Networking\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\",\"url\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\",\"name\":\"High-Performance GPU Server Hardware Topology and Cluster Networking - fibermall.com\",\"isPartOf\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#primaryimage\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#primaryimage\"},\"thumbnailUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png\",\"datePublished\":\"2024-09-27T09:26:42+00:00\",\"dateModified\":\"2024-09-27T09:26:45+00:00\",\"description\":\"NVSwitch is an NVIDIA switch chip encapsulated within the GPU module, not an independent external switch.\",\"breadcrumb\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#primaryimage\",\"url\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png\",\"contentUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png\",\"width\":1080,\"height\":546},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/www.fibermall.com\/blog.htm\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"High-Performance GPU Server Hardware Topology and Cluster Networking\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#website\",\"url\":\"https:\/\/www.fibermall.com\/blog.htm\/\",\"name\":\"fibermall.com\",\"description\":\"Optical Communication Expert\",\"publisher\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/www.fibermall.com\/blog.htm\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\",\"name\":\"fibermall.com\",\"url\":\"https:\/\/www.fibermall.com\/blog.htm\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg\",\"contentUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg\",\"width\":200,\"height\":67,\"caption\":\"fibermall.com\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/68e73044a5a439d9e8f42e16adcd0a86\",\"name\":\"FiberMall\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/4cf1580c4b121b9e7e65628346ff730750fa9bc63b9bddba66e390117d707ca3?s=96&d=mm&r=g\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/4cf1580c4b121b9e7e65628346ff730750fa9bc63b9bddba66e390117d707ca3?s=96&d=mm&r=g\",\"caption\":\"FiberMall\"},\"description\":\"One-stop supplier of professional optical communication products\",\"sameAs\":[\"https:\/\/www.fibermall.com\/blog\"],\"url\":\"https:\/\/www.fibermall.com\/blog.htm?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"High-Performance GPU Server Hardware Topology and Cluster Networking - fibermall.com","description":"NVSwitch is an NVIDIA switch chip encapsulated within the GPU module, not an independent external switch.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm","og_locale":"en_US","og_type":"article","og_title":"High-Performance GPU Server Hardware Topology and Cluster Networking","og_description":"Terminology and Basics Large model training typically utilizes single-machine, 8-GPU hosts to form clusters. The models include 8*{A100, A800, H100,","og_url":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm","og_site_name":"fibermall.com","article_published_time":"2024-09-27T09:26:42+00:00","article_modified_time":"2024-09-27T09:26:45+00:00","og_image":[{"width":1080,"height":546,"url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png","type":"image\/png"}],"author":"FiberMall","twitter_card":"summary_large_image","twitter_misc":{"Written by":"FiberMall","Est. reading time":"14 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#article","isPartOf":{"@id":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm"},"author":{"name":"FiberMall","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/68e73044a5a439d9e8f42e16adcd0a86"},"headline":"High-Performance GPU Server Hardware Topology and Cluster Networking","datePublished":"2024-09-27T09:26:42+00:00","dateModified":"2024-09-27T09:26:45+00:00","mainEntityOfPage":{"@id":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm"},"wordCount":2255,"commentCount":0,"publisher":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization"},"image":{"@id":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#primaryimage"},"thumbnailUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png","articleSection":["Blog","Networking"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm","url":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm","name":"High-Performance GPU Server Hardware Topology and Cluster Networking - fibermall.com","isPartOf":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#primaryimage"},"image":{"@id":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#primaryimage"},"thumbnailUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png","datePublished":"2024-09-27T09:26:42+00:00","dateModified":"2024-09-27T09:26:45+00:00","description":"NVSwitch is an NVIDIA switch chip encapsulated within the GPU module, not an independent external switch.","breadcrumb":{"@id":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#primaryimage","url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png","contentUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/09\/Host-Topology.png","width":1080,"height":546},{"@type":"BreadcrumbList","@id":"https:\/\/www.fibermall.com\/blog\/gpu-server-topology-and-networking.htm#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.fibermall.com\/blog.htm"},{"@type":"ListItem","position":2,"name":"High-Performance GPU Server Hardware Topology and Cluster Networking"}]},{"@type":"WebSite","@id":"https:\/\/www.fibermall.com\/blog.htm\/#website","url":"https:\/\/www.fibermall.com\/blog.htm\/","name":"fibermall.com","description":"Optical Communication Expert","publisher":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.fibermall.com\/blog.htm\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization","name":"fibermall.com","url":"https:\/\/www.fibermall.com\/blog.htm\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/","url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg","contentUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg","width":200,"height":67,"caption":"fibermall.com"},"image":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/68e73044a5a439d9e8f42e16adcd0a86","name":"FiberMall","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/4cf1580c4b121b9e7e65628346ff730750fa9bc63b9bddba66e390117d707ca3?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/4cf1580c4b121b9e7e65628346ff730750fa9bc63b9bddba66e390117d707ca3?s=96&d=mm&r=g","caption":"FiberMall"},"description":"One-stop supplier of professional optical communication products","sameAs":["https:\/\/www.fibermall.com\/blog"],"url":"https:\/\/www.fibermall.com\/blog.htm?author=1"}]}},"_links":{"self":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/13191","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=13191"}],"version-history":[{"count":3,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/13191\/revisions"}],"predecessor-version":[{"id":13218,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/13191\/revisions\/13218"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/media\/13200"}],"wp:attachment":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=13191"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=13191"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=13191"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}