Table of Contents
ToggleCore Port Ratio
Definition: Core port ratio refers to the proportional relationship between the ports on network devices (primarily switches) used to connect Computing Nodes (GPU Servers) and the ports used for Internal Network Interconnection when building GPU clusters (e.g., for AI training, HPC).
- Access/Leaf Ports: Ports used to directly connect GPU servers. Each GPU server typically connects to the network via multiple high-speed network cards (e.g., an 8-card server equipped with eight 200G NICs).
- Uplink/Spine/Trunk Ports: Ports used to connect to upper-layer switches or for switch-to-switch interconnection to form a larger-scale non-blocking network. In a Spine-Leaf architecture, these are the ports on Leaf switches connecting to Spine switches.
The Core Ratio Problem: The question is: Given a switch’s (usually a Leaf switch) total number of ports, how many should be allocated to servers and how many should be reserved for uplinks?
Common Ratio Examples:
- 1:1: For example, on a 128-port Leaf switch, 64 ports connect to servers and 64 ports connect up to Spines. This is the classic ratio for achieving a Full Bandwidth, Non-blocking network, ensuring sufficient vertical bandwidth for communication between any two servers.
- 2:1 or 3:1: For example, on a 128-port switch, 96 ports connect to servers and 32 connect up. This implies a certain degree of Oversubscription.
- Variable Ratios: In modern hyperscale clusters, ratios may be achieved through multiple layers and combinations of different port speeds to allow for more flexible designs.

Function:
- Determines Single-Point Bandwidth Limit: It directly determines the maximum bandwidth with which a single GPU server can simultaneously communicate with all other servers in the cluster.
- Balances Cost and Performance: A higher uplink port ratio (e.g., 1:1) offers optimal performance but results in the highest network equipment costs and port consumption. A lower ratio (e.g., 3:1) can save costs and support more servers under the same Leaf, but may create bottlenecks on the uplink.
- Influences Network Architecture Scale: The ratio, combined with the port count of Spine switches, determines the maximum number of servers the entire cluster can support.
Network Convergence
Contextual Definition:
- In general networking, Convergence usually refers to the time required for a network to recover from topology changes (like link failures or device reboots), recalculate/sync optimal paths, and reach a stable state. While fast convergence is vital for high availability, in the context of HPC and GPU Clusters, Network Convergence has a more specific, core meaning.
Network Convergence Ratio: This refers to the ratio of Total Downstream Bandwidth to Total Upstream or Peer Interconnect Bandwidth at a network aggregation point. It measures potential congestion points in the network.
Formula: Convergence Ratio = Total Access Bandwidth / Total Uplink Bandwidth
- Convergence Ratio = 1:1: Termed a “Non-blocking” or “No Oversubscription” network. At any given moment, if all access devices communicate at full bandwidth, no bottleneck will occur on the uplink. This is the gold standard for high-performance GPU clusters.
- Convergence Ratio > 1:1 (e.g., 3:1): Termed an “Oversubscription” network. This means if all downstream devices communicate at full load simultaneously, the upstream bandwidth is insufficient to carry the traffic, inevitably causing congestion and queuing. This value (e.g., 3:1) is the convergence ratio.

Function:
- Quantifies Network Congestion Risk: The convergence ratio is a key metric for measuring whether a network design meets application requirements. A low convergence ratio (close to 1:1) implies lower latency and higher predictability.
- Core Lever for Cost Control: Allowing a certain degree of oversubscription (e.g., 2:1) is the primary means of reducing network construction costs, as it serves more servers with fewer uplinks and core devices.
- Strong Correlation with Application Traffic Patterns: If the application traffic pattern is not “all-to-all” full bandwidth communication (e.g., primarily server-to-storage traffic, or computing tasks with communication intervals), a higher convergence ratio may be acceptable.
Interrelationships and Design Trade-offs
Means vs. Goal: Port ratio is the means to achieve a specific network convergence ratio, whereas the convergence ratio is one of the core goals the port ratio design aims to achieve.
Direct Determining Relationship: In a simple two-layer Spine-Leaf architecture, the port ratio of the Leaf switch directly determines the network convergence ratio of the first level (Leaf layer).
- If a Leaf switch adopts a 1:1 Port Ratio (half ports to servers, half ports up to Spines), then from the perspective of that Leaf switch, its Convergence Ratio is 1:1 (Non-blocking).
- If a 3:1 Port Ratio is adopted (assuming identical port speeds), then its Convergence Ratio is 3:1.
Multi-layer Accumulation: In large clusters, convergence can occur at multiple levels. For example, multiple GPUs inside a server interconnect via NVLink, then connect to a Leaf via NICs, then Leaf uplinks to Spine, and Spines may have further interconnections. The overall, end-to-end effective convergence ratio is the aggregate result of convergence at all levels. The design goal is usually to ensure the “Critical Path” (e.g., GPU-to-GPU communication path) has a convergence ratio of 1:1.
Deep Coupling with GPU Cluster Communication Patterns:
- All-Reduce Dominance: In modern AI training (especially Large Models), All-Reduce collective communication is the dominant pattern. This requires efficient paths between any two GPUs. If the network has a high convergence ratio, the uplink will easily become a bottleneck during the “many-to-many” communication phase of All-Reduce, severely slowing down training.
- Design Implications: Therefore, for GPU clusters designed for AI training, the Compute Plane Network (used for GPU-to-GPU communication) strongly pursues 1:1 Non-blocking Convergence. This is typically achieved through a 1:1 Port Ratio combined with network topologies like Clos/Fat-Tree. Storage planes and management networks, however, may use higher convergence ratios to save costs.

Summary Comparison Table
| Feature | Network Core Port Ratio | Network Convergence (Ratio) |
| Essence | Physical Design Decision: How switch ports are allocated and used. | Performance Metric/Concept: The proportional relationship between bandwidth supply and demand. |
| Focus | The division and connection methods of device port resources. | The degree of overload or redundancy in network link capacity. |
| Manifestation | “This Leaf switch has 64 server ports and 32 uplink ports.” | “From server aggregation to Spine, the network convergence ratio is 2:1.” |
| Function | Implements network topology; balances single-point access density with uplink capacity. | Quantifies potential network performance bottlenecks; key for balancing cost vs. performance. |
| Relationship | Port ratio is the physical means to achieve the target convergence ratio. In a single layer, the ratio values are numerically equal. | Convergence ratio is the core standard for evaluating whether the port ratio design meets application requirements. |
Conclusion: In GPU cluster network design, one must determine the target convergence ratio (usually pursuing 1:1 non-blocking) based on upper-layer applications (specifically AI training jobs), and then achieve this goal through meticulous port ratio and topology design (such as Fat-Tree, Dragonfly+). Understanding the concepts of “Port Ratio” and “Network Convergence” and their relationship is the foundation for designing or evaluating a high-performance GPU cluster network. Simply put, port ratio is “how to do it,” and convergence ratio is “why to do it this way and what the effect is.”
Related Products:
-
NVIDIA NVIDIA(Mellanox) MCX653106A-ECAT-SP ConnectX-6 InfiniBand/VPI Adapter Card, HDR100/EDR/100G, Dual-Port QSFP56, PCIe3.0/4.0 x16, Tall Bracket
$1100.00
-
NVIDIA MCX623106AN-CDAT SmartNIC ConnectX®-6 Dx EN Network Interface Card, 100GbE Dual-Port QSFP56, PCIe4.0 x 16, Tall&Short Bracket
$1200.00
-
NVIDIA NVIDIA(Mellanox) MCX653105A-ECAT-SP ConnectX-6 InfiniBand/VPI Adapter Card, HDR100/EDR/100G, Single-Port QSFP56, PCIe3.0/4.0 x16, Tall bracket
$965.00
-
NVIDIA NVIDIA(Mellanox) MCX516A-CCAT SmartNIC ConnectX®-5 EN Network Interface Card, 100GbE Dual-Port QSFP28, PCIe3.0 x 16, Tall&Short Bracket
$985.00
-
NVIDIA NVIDIA(Mellanox) MCX653105A-HDAT-SP ConnectX-6 InfiniBand/VPI Adapter Card, HDR/200GbE, Single-Port QSFP56, PCIe3.0/4.0 x16, Tall Bracket
$1400.00
-
NVIDIA NVIDIA(Mellanox) MCX653106A-HDAT-SP ConnectX-6 InfiniBand/VPI Adapter Card, HDR/200GbE, Dual-Port QSFP56, PCIe3.0/4.0 x16, Tall Bracket
$1600.00
-
NVIDIA B3140H BlueField-3 8 Arm-Cores SuperNIC, E-series HHHL, 400GbE (Default Mode)/NDR IB, Single-port QSFP112, PCle Gen5.0 x16, 16GB Onboard DDR, Integrated BMC, Crypto Disabled
$4390.00
-
NVIDIA NVIDIA(Mellanox) MCX75310AAS-NEAT ConnectX-7 InfiniBand/VPI Adapter Card, NDR/400G, Single-port OSFP, PCIe 5.0x 16, Tall Bracket
$2200.00
-
NVIDIA NVIDIA(Mellanox) MCX75510AAS-NEAT ConnectX-7 InfiniBand/VPI Adapter Card, NDR/400G, Single-port OSFP, PCIe 5.0x 16, Tall Bracket
$1650.00
-
NVIDIA ConnectX-8 C8180 (900-9X81E-00EX-ST0) HHHL SuperNIC, 800Gbs XDR IB & Ethernet, Single-cage OSFP, PCIe 6 x16 with x16 PCIe Socket Direct,Tall bracket
$1980.00
