{"id":8113,"date":"2024-05-14T05:53:38","date_gmt":"2024-05-14T05:53:38","guid":{"rendered":"https:\/\/www.fibermall.com\/blog\/?p=8113"},"modified":"2024-05-14T05:53:44","modified_gmt":"2024-05-14T05:53:44","slug":"a100-h100-gh200-cluster","status":"publish","type":"post","link":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm","title":{"rendered":"A100\/H100\/GH200 Cluster: Network Architecture and Optical Module Requirements"},"content":{"rendered":"\n<p>Traditional data centers have undergone a transition from a three-tier architecture to a leaf-spine architecture, primarily to accommodate the growth of east-west traffic within the data center. As the process of data migration to the cloud continues to accelerate, the scale of cloud computing data centers continues to expand. Applications such as virtualization and hyper-converged systems adopted in these data centers have driven a significant increase in east-west traffic\u2014according to previous data from Cisco, in 2021, internal data center traffic accounted for over 70% of data center-related traffic.<\/p>\n\n\n\n<p>Taking the transition from traditional three-tier architecture to leaf-spine architecture as an example, the number of optical modules required in a leaf-spine network architecture can increase by up to tens of times.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/white-paper-c11-737022_1.jpg\" alt=\"white-paper-c11-737022_1\" class=\"wp-image-8115\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/white-paper-c11-737022_1.jpg 616w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/white-paper-c11-737022_1-300x185.jpg 300w\" sizes=\"(max-width: 616px) 100vw, 616px\" \/><\/figure>\n\n\n\n<p>Network Architecture Requirements for Large-Scale AI Clusters<\/p>\n\n\n\n<p>Considering the need to alleviate network bottlenecks, the network architecture for large-scale AI clusters must meet the requirements of high bandwidth, low latency, and lossless transmission. AI computing centers generally adopt a Fat-Tree network architecture, which features a non-blocking network. Additionally, to avoid inter-node interconnect bottlenecks, NVIDIA employs NVLink to enable efficient inter-GPU communication. Compared to PCIe, NVLink offers higher bandwidth advantages, serving as the foundation for NVIDIA&#8217;s shared memory architecture and creating a new demand for optical interconnects between GPUs.<\/p>\n\n\n\n<p>A100 Network Structure and Optical Module Requirements<\/p>\n\n\n\n<p>The basic deployment structure for each DGX A100 SuperPOD consists of 140 servers (each server with 8 GPUs) and switches (each switch with 40 ports, each port at 200G). The network topology is an InfiniBand (IB) Fat-Tree structure. Regarding the number of network layers, a three-layer network structure (server-leaf switch-spine switch-core switch) is deployed for 140 servers, with the corresponding number of cables for each layer being 1120-1124-1120, respectively. Assuming copper cables are used between servers and switches, and based on one cable corresponding to two 200G optical modules, the ratio of GPU:switch:optical module is 1:0.15:4. If an all-optical network is used, the ratio becomes GPU:switch:optical module = 1:0.15:6.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/developer_c087f74-1024x570.png\" alt=\"developer_c087f74\" class=\"wp-image-8116\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/developer_c087f74-1024x570.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/developer_c087f74-300x167.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/developer_c087f74-768x427.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/developer_c087f74.png 1148w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/5ZCez_5CQB3B-1024x477.png\" alt=\"5ZCez_5CQB3B\" class=\"wp-image-8117\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/5ZCez_5CQB3B-1024x477.png 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/5ZCez_5CQB3B-300x140.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/5ZCez_5CQB3B-768x358.png 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/5ZCez_5CQB3B.png 1280w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>H100 Network Structure and Optical Module Requirements<\/p>\n\n\n\n<p>The basic deployment structure for each DGX H100 SuperPOD consists of 32 servers (each server with 8 GPUs) and 12 switches. The network topology is an IB Fat-Tree structure, with each switch port operating at 400G and capable of being combined into an 800G port. For a 4SU cluster, assuming an all-optical network and a three-layer Fat-Tree architecture, <a href=\"https:\/\/www.fibermall.com\/store-21995-400g-ndr-infiniband.htm\" target=\"_blank\" rel=\"noreferrer noopener\">400G optical modules <\/a>are used between servers and leaf switches, while 800G optical modules are used between leaf-spine and spine-core switches. The number of 400G optical modules required is 3284=256, and the number of 800G optical modules is 3282.5=640. Therefore, the ratio of GPU:switch:400G optical module:800G optical module is 1:0.08:1:2.5.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/rzcF8_94mA6j-1024x416.jpg\" alt=\"rzcF8_94mA6j\" class=\"wp-image-8118\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/rzcF8_94mA6j-1024x416.jpg 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/rzcF8_94mA6j-300x122.jpg 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/rzcF8_94mA6j-768x312.jpg 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/rzcF8_94mA6j-1536x624.jpg 1536w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/rzcF8_94mA6j.jpg 1600w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>For a single GH200 cluster, which consists of 256 interconnected super-chip GPUs using a two-tier fat-tree network structure, both tiers are built with NVLink switches. The first tier (between servers and Level 1 switches) uses 96 switches, while Level 2 employs 36 switches. Each NVLink switch has 32 ports, with each port having a speed of 800G. Given that the NVLink 4.0\u2019s bidirectional aggregated bandwidth is 900GB\/s, and unidirectional is 450GB\/s, the total uplink bandwidth for the access layer in a 256-card cluster is 115,200GB\/s. Considering the fat-tree architecture and the 800G optical module transmission rate (100GB\/s), the total requirement for 800G optical modules is 2,304 units. Therefore, within the GH200 cluster, the ratio of GPUs to optical modules is 1:9. When interconnecting multiple GH200 clusters, referencing the H100 architecture, under a three-tier network structure, the demand for GPUs to 800G optical modules is 1:2.5; under a two-tier network, it is 1:1.5. Thus, when interconnecting multiple GH200s, the upper limit for the GPU to 800G optical module ratio is 1:(9+2.5) = 1:11.5.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/nvidia-grace-hopper-gh200-nvlink-fabric.jpg\" alt=\"nvidia-grace-hopper-gh200-nvlink-fabric\" class=\"wp-image-8119\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/nvidia-grace-hopper-gh200-nvlink-fabric.jpg 1015w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/nvidia-grace-hopper-gh200-nvlink-fabric-300x111.jpg 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/nvidia-grace-hopper-gh200-nvlink-fabric-768x285.jpg 768w\" sizes=\"(max-width: 1015px) 100vw, 1015px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/NVIDIA-GH-Superchip-System.png\" alt=\"NVIDIA GH Superchip System\" class=\"wp-image-8120\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/NVIDIA-GH-Superchip-System.png 640w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/NVIDIA-GH-Superchip-System-300x147.png 300w\" sizes=\"(max-width: 640px) 100vw, 640px\" \/><\/figure>\n\n\n\n<p>In summary, as computational clusters continue to enhance network performance, the demand for high-speed optical modules becomes more elastic. Taking NVIDIA clusters as an example, the network card interface rate adapted by the accelerator card is closely related to its network protocol bandwidth. The A100 GPU supports PCIe 4.0, with a maximum unidirectional bandwidth of 252Gb\/s, hence the PCIe network card rate must be less than 252Gb\/s, pairing with Mellanox HDR 200Gb\/s Infiniband network cards. The H100 GPU supports PCIe 5.0, with a maximum unidirectional bandwidth of 504Gb\/s, thus pairing with Mellanox NDR 400Gb\/s Infiniband network cards. Therefore, upgrading from A100 to H100, the corresponding optical module demand increases from 200G to 800G (two 400G ports combined into one 800G); while the GH200 uses NVLink for inter-card connectivity, with unidirectional bandwidth increased to 450GB\/s, further increasing the elasticity for the 800G demand. Suppose the H100 cluster upgrades from PCIe 5.0 to PCIe 6.0, with the maximum unidirectional bandwidth increased to 1024Gb\/s. In that case, the access layer network card rate can be raised to 800G, meaning the access layer can use 800G optical modules, and the demand elasticity for a single card corresponding to 800G optical modules in the cluster would double.<\/p>\n\n\n\n<p>Meta\u2019s computational cluster architecture and application previously released the \u201cResearch SuperCluster\u201d&nbsp;project for training the LLaMA model. In the second phase of the RSC project, Meta deployed a total of 2,000 A100 servers, containing 16,000 A100 GPUs. The cluster includes 2,000 switches and 48,000 links, corresponding to a three-tier CLOS network architecture. If a full optical network is adopted, it corresponds to 96,000 200G optical modules, meaning the ratio of A100 GPUs to optical modules is 1:6, consistent with the previously calculated A100 architecture.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology.jpg\" alt=\"meta-networking-scale-32k-scale-topology\" class=\"wp-image-8121\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology.jpg 529w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-300x157.jpg 300w\" sizes=\"(max-width: 529px) 100vw, 529px\" \/><\/figure>\n\n\n\n<p>Meta has implemented a training infrastructure for LLaMA3 using H100 GPUs, which includes clusters with both InfiniBand and Ethernet, capable of supporting up to 32,000 GPUs. For the Ethernet solution, according to information disclosed by Meta, the computing cluster still employs a converged leaf-spine network architecture. Each rack contains 2 servers connected to 1 Top-of-Rack (TOR) switch (using Wedge 400), with a total of 252 servers in a cluster. The cluster switches use Minipack2 OCP rack switches, with 18 cluster switches in total, resulting in a convergence ratio of 3.5:1. There are 18 aggregation layer switches (using Arista 7800R3), with a convergence ratio of 7:1. The cluster primarily uses 400G optical modules. From the cluster architecture perspective, the Ethernet solution still requires further breakthroughs at the protocol level to promote the construction of a non-blocking network, with attention to the progress of organizations such as the Ethernet Alliance.<\/p>\n\n\n\n<p>AWS has launched the second generation of EC2 Ultra Clusters, which include the H100 GPU and their proprietary Trainium ASIC solution. The AWS EC2 Ultra Clusters P5 instances (i.e., the H100 solution) provide an aggregate network bandwidth of 3200 Gbps and support GPUDirect RDMA, with a maximum networking capacity of 20,000 GPUs. The Trn1n instances (proprietary Trainium solution) feature a 16-card cluster providing 1600 Gbps of aggregate network bandwidth, supporting up to 30,000 ASICs networked, corresponding to 6 EFlops of computing power.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/aws-ec2-ultrascluster-block-diagram-1024x494.jpg\" alt=\"aws-ec2-ultrascluster-block-diagram\" class=\"wp-image-8122\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/aws-ec2-ultrascluster-block-diagram-1024x494.jpg 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/aws-ec2-ultrascluster-block-diagram-300x145.jpg 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/aws-ec2-ultrascluster-block-diagram-768x370.jpg 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/aws-ec2-ultrascluster-block-diagram-1536x740.jpg 1536w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/aws-ec2-ultrascluster-block-diagram.jpg 1745w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/image-2-1024x544.jpg\" alt=\"image-2\" class=\"wp-image-8123\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/image-2-1024x544.jpg 1024w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/image-2-300x159.jpg 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/image-2-768x408.jpg 768w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/image-2-1536x815.jpg 1536w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/image-2.jpg 2046w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p>The interconnection between AWS EC2 Ultra Clusters cards uses NVLink (for the H100 solution) and NeuronLink (for the Trainium solution), with cluster interconnection using their proprietary EFA network adapter. Compared to Nvidia\u2019s solution, AWS\u2019s proprietary Trainium ASIC cluster has an estimated uplink bandwidth of 100G per card (1600G aggregate bandwidth \/ 16 cards = 100G), hence there is currently no demand for <a href=\"https:\/\/www.fibermall.com\/store-21994-800g-ndr-infiniband.htm\" target=\"_blank\" rel=\"noreferrer noopener\">800G<\/a> optical modules in AWS\u2019s architecture.<\/p>\n\n\n\n<p>Google\u2019s latest computing cluster is composed of TPU arrays configured in a three-dimensional torus. A one-dimensional torus corresponds to each TPU connected to two adjacent TPUs, a two-dimensional torus consists of two orthogonal rings, corresponding to each TPU connected to four adjacent TPUs; Google\u2019s TPUv4 represents a three-dimensional torus, with each TPU connected to six adjacent TPUs.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/Google-Machine-Learning-Supercomputer-With-An-Optically-Reconfigurable-Interconnect-_Page_11-746x420-1.jpg\" alt=\"Google-Machine-Learning-Supercomputer-With-An-Optically-Reconfigurable-Interconnect-_Page_11-746x420\" class=\"wp-image-8124\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/Google-Machine-Learning-Supercomputer-With-An-Optically-Reconfigurable-Interconnect-_Page_11-746x420-1.jpg 746w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/Google-Machine-Learning-Supercomputer-With-An-Optically-Reconfigurable-Interconnect-_Page_11-746x420-1-300x169.jpg 300w\" sizes=\"(max-width: 746px) 100vw, 746px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/Bidirectional-CWDM4-optical-transceiver.jpg\" alt=\"Bidirectional CWDM4 optical transceiver\" class=\"wp-image-8125\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/Bidirectional-CWDM4-optical-transceiver.jpg 664w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/Bidirectional-CWDM4-optical-transceiver-300x172.jpg 300w\" sizes=\"(max-width: 664px) 100vw, 664px\" \/><\/figure>\n\n\n\n<p>Based on this, a 3D network structure of 444=64 TPUs is constructed within each cabinet. The external part of the 3D structure connects to the OCS, with an interconnection of 4096 TPUs corresponding to 64 cabinets and 48 OCS switches, which equals 48*64=6144 optical modules. Internally, DAC connections are used (18000 cables), resulting in a TPU to optical module ratio of 1:1.5. Under the OCS solution, the optical modules need to adopt a wavelength-division multiplexing solution and add circulators to reduce the number of fibers, with the optical module solution having customized features (800G VFR8).<\/p>\n<style>\r\n\r\n        .lwrp.link-whisper-related-posts{\r\n            \r\n            margin-top: 40px;\nmargin-bottom: 30px;\r\n        }\r\n        .lwrp .lwrp-title{\r\n            \r\n            \r\n        }\r\n        .lwrp .lwrp-description{\r\n            \r\n            \r\n\r\n        }\r\n        .lwrp .lwrp-list-container{\r\n        }\r\n        .lwrp .lwrp-list-multi-container{\r\n            display: flex;\r\n        }\r\n        .lwrp .lwrp-list-double{\r\n            width: 48%;\r\n        }\r\n        .lwrp .lwrp-list-triple{\r\n            width: 32%;\r\n        }\r\n        .lwrp .lwrp-list-row-container{\r\n            display: flex;\r\n            justify-content: space-between;\r\n        }\r\n        .lwrp .lwrp-list-row-container .lwrp-list-item{\r\n            width: calc(100% - 20px);\r\n        }\r\n        .lwrp .lwrp-list-item:not(.lwrp-no-posts-message-item){\r\n            \r\n            list-style: decimal;\r\n        }\r\n        .lwrp .lwrp-list-item img{\r\n            max-width: 100%;\r\n            height: auto;\r\n        }\r\n        .lwrp .lwrp-list-item.lwrp-empty-list-item{\r\n            background: initial !important;\r\n        }\r\n        .lwrp .lwrp-list-item .lwrp-list-link .lwrp-list-link-title-text,\r\n        .lwrp .lwrp-list-item .lwrp-list-no-posts-message{\r\n            \r\n                \r\n        }\r\n        @media screen and (max-width: 480px) {\r\n            .lwrp.link-whisper-related-posts{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-title{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-description{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-list-multi-container{\r\n                flex-direction: column;\r\n            }\r\n            .lwrp .lwrp-list-multi-container ul.lwrp-list{\r\n                margin-top: 0px;\r\n                margin-bottom: 0px;\r\n                padding-top: 0px;\r\n                padding-bottom: 0px;\r\n            }\r\n            .lwrp .lwrp-list-double,\r\n            .lwrp .lwrp-list-triple{\r\n                width: 100%;\r\n            }\r\n            .lwrp .lwrp-list-row-container{\r\n                justify-content: initial;\r\n                flex-direction: column;\r\n            }\r\n            .lwrp .lwrp-list-row-container .lwrp-list-item{\r\n                width: 100%;\r\n            }\r\n            .lwrp .lwrp-list-item:not(.lwrp-no-posts-message-item){\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-list-item .lwrp-list-link .lwrp-list-link-title-text,\r\n            .lwrp .lwrp-list-item .lwrp-list-no-posts-message{\r\n                \r\n                    \r\n            }\r\n        }<\/style>\r\n<div id=\"link-whisper-related-posts-widget\" class=\"link-whisper-related-posts lwrp\">\r\n            <h3 class=\"lwrp-title\">Related Posts<\/h3>    \r\n        <div class=\"lwrp-list-container\">\r\n                                            <ul class=\"lwrp-list lwrp-list-single\">\r\n                    <li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/internet-switch-hub.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">Ethernet Switch Hub: The Ultimate Guide to Choosing the Right Device<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/10gb-ethernet-switch.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">Unlock Lightning-Fast Speeds: Your Ultimate Guide to 10GB Ethernet Switches<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/fiber-in-datacenter-connect-with-mpo.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">How Backbone Fiber in a Data Center is Connected with MPO Connector?<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/400g-dwdm-combine-qsfp-dd-with-dwdm-coherent.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">400G DWDM: Combine QSFP-DD Transceiver with DWDM Coherent<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/24-port-poe-switch.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">Unlock the Potential of Your Network with a 24-Port PoE Switch<\/span><\/a><\/li>                <\/ul>\r\n                        <\/div>\r\n<\/div>","protected":false},"excerpt":{"rendered":"<p>Traditional data centers have undergone a transition from a three-tier architecture to a leaf-spine architecture, primarily to accommodate the growth of east-west traffic within the data center. As the process of data migration to the cloud continues to accelerate, the scale of cloud computing data centers continues to expand. Applications such as virtualization and hyper-converged [&hellip;]<\/p>\n","protected":false},"author":5,"featured_media":8129,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"set","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":"","_wpscppro_dont_share_socialmedia":false,"_wpscppro_custom_social_share_image":0,"_facebook_share_type":"","_twitter_share_type":"","_linkedin_share_type":"","_pinterest_share_type":"","_linkedin_share_type_page":"","_instagram_share_type":"","_medium_share_type":"","_threads_share_type":"","_google_business_share_type":"","_selected_social_profile":[],"_wpsp_enable_custom_social_template":false,"_wpsp_social_scheduling":{"enabled":false,"datetime":null,"platforms":[],"status":"template_only","dateOption":"today","timeOption":"now","customDays":"","customHours":"","customDate":"","customTime":"","schedulingType":"absolute"},"_wpsp_active_default_template":true},"categories":[2,29],"tags":[],"class_list":["post-8113","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog","category-networking"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v20.13 (Yoast SEO v25.8) - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>A100\/H100\/GH200 Cluster: Network Architecture | FiberMall<\/title>\n<meta name=\"description\" content=\"The basic deployment structure for each DGX H100 SuperPOD consists of 32 servers (each server with 8 GPUs) and 12 switches.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"A100\/H100\/GH200 Cluster: Network Architecture and Optical Module Requirements\" \/>\n<meta property=\"og:description\" content=\"Traditional data centers have undergone a transition from a three-tier architecture to a leaf-spine architecture, primarily to accommodate the growth of\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm\" \/>\n<meta property=\"og:site_name\" content=\"fibermall.com\" \/>\n<meta property=\"article:published_time\" content=\"2024-05-14T05:53:38+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2024-05-14T05:53:44+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-1.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"529\" \/>\n\t<meta property=\"og:image:height\" content=\"352\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Brian\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Brian\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"9 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#article\",\"isPartOf\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm\"},\"author\":{\"name\":\"Brian\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/9421292f8817d668e9e64f94dfe9b901\"},\"headline\":\"A100\/H100\/GH200 Cluster: Network Architecture and Optical Module Requirements\",\"datePublished\":\"2024-05-14T05:53:38+00:00\",\"dateModified\":\"2024-05-14T05:53:44+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm\"},\"wordCount\":1325,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#primaryimage\"},\"thumbnailUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-1.jpg\",\"articleSection\":[\"Blog\",\"Networking\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm\",\"url\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm\",\"name\":\"A100\/H100\/GH200 Cluster: Network Architecture | FiberMall\",\"isPartOf\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#primaryimage\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#primaryimage\"},\"thumbnailUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-1.jpg\",\"datePublished\":\"2024-05-14T05:53:38+00:00\",\"dateModified\":\"2024-05-14T05:53:44+00:00\",\"description\":\"The basic deployment structure for each DGX H100 SuperPOD consists of 32 servers (each server with 8 GPUs) and 12 switches.\",\"breadcrumb\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#primaryimage\",\"url\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-1.jpg\",\"contentUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-1.jpg\",\"width\":529,\"height\":352,\"caption\":\"meta-networking-scale-32k-scale-topology\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/www.fibermall.com\/blog.htm\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"A100\/H100\/GH200 Cluster: Network Architecture and Optical Module Requirements\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#website\",\"url\":\"https:\/\/www.fibermall.com\/blog.htm\/\",\"name\":\"fibermall.com\",\"description\":\"Optical Communication Expert\",\"publisher\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/www.fibermall.com\/blog.htm\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\",\"name\":\"fibermall.com\",\"url\":\"https:\/\/www.fibermall.com\/blog.htm\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg\",\"contentUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg\",\"width\":200,\"height\":67,\"caption\":\"fibermall.com\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/9421292f8817d668e9e64f94dfe9b901\",\"name\":\"Brian\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/c196835c5ddcc99559fcfaa0046ffbd895891b78d40ab46bf34f7a5ecd73dfd9?s=96&d=mm&r=g\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/c196835c5ddcc99559fcfaa0046ffbd895891b78d40ab46bf34f7a5ecd73dfd9?s=96&d=mm&r=g\",\"caption\":\"Brian\"},\"description\":\"Optical Network Engineer\",\"sameAs\":[\"https:\/\/www.fibermall.com\"],\"url\":\"https:\/\/www.fibermall.com\/blog.htm?author=5\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"A100\/H100\/GH200 Cluster: Network Architecture | FiberMall","description":"The basic deployment structure for each DGX H100 SuperPOD consists of 32 servers (each server with 8 GPUs) and 12 switches.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm","og_locale":"en_US","og_type":"article","og_title":"A100\/H100\/GH200 Cluster: Network Architecture and Optical Module Requirements","og_description":"Traditional data centers have undergone a transition from a three-tier architecture to a leaf-spine architecture, primarily to accommodate the growth of","og_url":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm","og_site_name":"fibermall.com","article_published_time":"2024-05-14T05:53:38+00:00","article_modified_time":"2024-05-14T05:53:44+00:00","og_image":[{"width":529,"height":352,"url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-1.jpg","type":"image\/jpeg"}],"author":"Brian","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Brian","Est. reading time":"9 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#article","isPartOf":{"@id":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm"},"author":{"name":"Brian","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/9421292f8817d668e9e64f94dfe9b901"},"headline":"A100\/H100\/GH200 Cluster: Network Architecture and Optical Module Requirements","datePublished":"2024-05-14T05:53:38+00:00","dateModified":"2024-05-14T05:53:44+00:00","mainEntityOfPage":{"@id":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm"},"wordCount":1325,"commentCount":0,"publisher":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization"},"image":{"@id":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#primaryimage"},"thumbnailUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-1.jpg","articleSection":["Blog","Networking"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm","url":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm","name":"A100\/H100\/GH200 Cluster: Network Architecture | FiberMall","isPartOf":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#primaryimage"},"image":{"@id":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#primaryimage"},"thumbnailUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-1.jpg","datePublished":"2024-05-14T05:53:38+00:00","dateModified":"2024-05-14T05:53:44+00:00","description":"The basic deployment structure for each DGX H100 SuperPOD consists of 32 servers (each server with 8 GPUs) and 12 switches.","breadcrumb":{"@id":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#primaryimage","url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-1.jpg","contentUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/meta-networking-scale-32k-scale-topology-1.jpg","width":529,"height":352,"caption":"meta-networking-scale-32k-scale-topology"},{"@type":"BreadcrumbList","@id":"https:\/\/www.fibermall.com\/blog\/a100-h100-gh200-cluster.htm#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.fibermall.com\/blog.htm"},{"@type":"ListItem","position":2,"name":"A100\/H100\/GH200 Cluster: Network Architecture and Optical Module Requirements"}]},{"@type":"WebSite","@id":"https:\/\/www.fibermall.com\/blog.htm\/#website","url":"https:\/\/www.fibermall.com\/blog.htm\/","name":"fibermall.com","description":"Optical Communication Expert","publisher":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.fibermall.com\/blog.htm\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization","name":"fibermall.com","url":"https:\/\/www.fibermall.com\/blog.htm\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/","url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg","contentUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg","width":200,"height":67,"caption":"fibermall.com"},"image":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/9421292f8817d668e9e64f94dfe9b901","name":"Brian","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/c196835c5ddcc99559fcfaa0046ffbd895891b78d40ab46bf34f7a5ecd73dfd9?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/c196835c5ddcc99559fcfaa0046ffbd895891b78d40ab46bf34f7a5ecd73dfd9?s=96&d=mm&r=g","caption":"Brian"},"description":"Optical Network Engineer","sameAs":["https:\/\/www.fibermall.com"],"url":"https:\/\/www.fibermall.com\/blog.htm?author=5"}]}},"_links":{"self":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/8113","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/users\/5"}],"replies":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=8113"}],"version-history":[{"count":4,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/8113\/revisions"}],"predecessor-version":[{"id":8132,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/8113\/revisions\/8132"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/media\/8129"}],"wp:attachment":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=8113"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=8113"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=8113"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}