Sarah examined the two quotes displayed on her monitor. The InfiniBand solution for her team’s new AI cluster would cost $680,000. The RoCEv2 Ethernet option using QSFP-DD 800G came in at $410,000. The performance specs showed only a 5% difference in all-reduce benchmarks.
“That’s $270,000 for 5%,” she thought. Her CFO would ask hard questions about that trade-off.
She discovered engineers who operated 512-GPU clusters on RoCEv2 through her exploration of Reddit threads. PFC and ECN with buffer settings needed a precise configuration to achieve successful operation. Get those right, and Ethernet delivered 90% of InfiniBand’s performance at 60% of the cost.
Sarah selected RoCEv2 through QSFP-DD. The cluster started operations after six months. The system processed training jobs through its entire operational period. The organization used their $270,000 savings to double its GPU capacity.
Sarah’s dilemma arises when organizations build AI infrastructure. The first stage requires organizations to choose between multiple form factor options. The second stage requires organizations to make their protocol selection. The third stage requires organizations to determine their size requirements. The guide establishes a framework for QSFP-DD AI deployment through actual bandwidth measurements, form factor evaluations, and deployment procedures.

What you’ll learn: GPU-to-bandwidth sizing formulas, QSFP-DD vs OSFP for AI clusters, RoCEv2 vs InfiniBand trade-offs, and module selection for different AI workloads.
New to QSFP-DD? Start with our complete QSFP-DD guide for fundamentals.
Table of Contents
ToggleAI Workload Bandwidth Requirements
Each H100 requires 400G of scale-out bandwidth. Do the math.
The training process of contemporary artificial intelligence systems requires multiple graphics processing units. The training of large language models needs distributed systems that utilize hundreds or thousands of graphics processing units. The training process requires each graphics processing unit to maintain constant communication with other units to share gradients, synchronize weights, and manage batch operations.
The NVLink system operates exclusively within a server to establish GPU-to-GPU connections. The discussion focuses on the scale-out network, which enables interconnection between multiple servers. A single NVIDIA H100 pushes 400G of Ethernet/InfiniBand traffic for this scale-out communication. H200 maintains the same 400G. The upcoming B200 and B300 generations push 800G per GPU.
GPU Bandwidth by Generation
| GPU | Scale-Out Bandwidth | Typical Use Case |
| H100 | 400G | Current training clusters |
| H200 | 400G | Memory-intensive training |
| B200 | 800G | Next-gen training |
| B300 | 800G+ | Large-scale AI factories |
Source: NVIDIA technical specifications
Cluster Sizing Formula
The practitioners have established a rule that states that 800G becomes essential when there are 128 or more GPUs in use.
The math is straightforward. At 128 H100 GPUs, you need 51.2T of aggregate bandwidth for non-blocking connectivity. The required bandwidth equals 128 times 400G. Your spine layer needs to handle this aggregation.
With 400G spine connections, you’d need 128 spine ports just for the GPU fabric. With 800G QSFP-DD, you cut that to 64 ports. The switch count decreases by half. The expenses decrease. The system becomes less complicated.
For context: A 1,024-GPU cluster needs 400T of aggregate bandwidth at 400G per GPU. At 800G spine speeds, you need 512 × 800G spine ports. The mathematical calculations become unending as clusters grow.

Sizing guidelines:
- Under 128 GPUs: 400G spine is fine
- 128-512 GPUs: 800G spine recommended
- 512+ GPUs: 800G spine essential, 1.6T on horizon
Reddit insight: “General rule—800G spine at 128+ GPUs. 400G works fine below that.”
QSFP-DD vs OSFP for AI Clusters
The form factor decision that trips up AI architects.
NVIDIA’s latest DGX systems ship with OSFP. But your existing data center probably runs QSFP-DD for 400G. Which do you choose for your AI cluster?
Decision Matrix for AI
| Factor | QSFP-DD | OSFP |
| Backward compatibility | 400G QSFP-DD | None (OSFP only) |
| Best for | Mixed training/inference | Dedicated training clusters |
| Thermal handling | Good (up to 18W) | Better (up to 20W+) |
| NVIDIA alignment | Moderate | Strong (DGX default) |
| Flexibility | High | Purpose-built |

When to Choose QSFP-DD for AI
Choose QSFP-DD when:
- You have an existing 400G QSFP-DD infrastructure to leverage
- You’re running mixed workloads (training and inference)
- You need flexibility to reuse switches between different environments
- Budget constraints favor backward compatibility
- You’re building general-purpose infrastructure, not dedicated AI training
Real example: A financial services firm built a 256-GPU cluster using QSFP-DD 800G. They connected it to their existing 400G storage network without adapters or separate switches. The backward compatibility saved $180,000 in additional hardware.
When OSFP Makes Sense
Choose OSFP when:
- You’re building a dedicated AI training cluster from scratch
- Maximum thermal headroom matters (20W+ modules)
- You’re deploying NVIDIA DGX systems exclusively
- You want alignment with NVIDIA’s roadmap
- You don’t need backward compatibility with 400G
Reddit insight: “QSFP-DD if you want to reuse switches between training and inference. OSFP if you’re building dedicated training clusters.”
AI Network Architecture
Frontend vs backend—the AI networking split.
AI clusters operate through three separate networks which have different requirements for bandwidth and latency, and QSFP-DD module selection.
Marcus experienced this lesson through difficult experiences. His team constructed a stunning 512-GPU cluster, which used 800G QSFP-DD spine connections. The team correctly selected all necessary modules to build the compute fabric. The team then found out that their storage network did not provide enough data to operate the GPUs at full capacity. For guidance on cable selection for AI clusters, see our QSFP-DD cable compatibility guide.
The computer network operated without any interruptions. The storage network experienced a 4:1 oversubscription problem. The GPUs remained inactive because they required data. The cluster managed to achieve 50 percent of its maximum performance capacity.
Marcus needed to create a new storage network design that used special QSFP-DD 400G connections. The project experienced a three-month delay. The lesson: AI networking requires more than GPU-to-GPU bandwidth connections.

Backend Network (GPU Fabric)
The backend network carries training traffic—gradient exchange, parameter synchronization, and all-reduce operations. This is your critical path.
Architecture: Two-tier spine-leaf with full non-blocking connectivity
Speed: 800G QSFP-DD spine, 400G or 800G to GPUs
Protocol: RoCEv2 or InfiniBand
Latency requirement: Under 2 microseconds switch-to-switch
The backend network requires operation without any interruptions. The whole training job stops when any packet gets lost. The Ethernet network requires both Priority Flow Control (PFC) and Explicit Congestion Notification (ECN) as essential components.
Storage Network
AI training reads massive datasets—terabytes per hour. The storage network feeds this data hunger.
Architecture: Separate fabric or converged with backend
Speed: 400G QSFP-DD is typically sufficient
Oversubscription: 2:1 or 4:1 common (unlike backend’s 1:1)
Protocol: NFS, GPFS, or object storage protocols
Storage networks can tolerate oversubscription because data access patterns are bursty, not constant like GPU communication. But don’t oversubscribe too heavily—starved GPUs waste money.
Frontend/Out-of-Band Network
Management, monitoring, and external connectivity.
Architecture: Traditional enterprise network
Speed: 100G or 200G QSFP-DD
Purpose: Cluster management, job scheduling, user access
The frontend network doesn’t need QSFP-DD 800G. 100G or 200G handles management traffic fine. Save your 800G ports for where they matter—the backend fabric.
Protocol Selection: RoCEv2 vs InfiniBand
90% of performance at 60% of cost.
The protocol debate divides AI infrastructure teams. InfiniBand has maintained its status as the primary technology for high-performance computing during multiple decades. RoCEv2 (RDMA over Converged Ethernet) provides an economic solution that competes with that market dominance.

Performance Comparison
| Metric | InfiniBand NDR | RoCEv2 800G | Difference |
| Latency | ~600ns | ~800ns | 33% higher |
| All-reduce throughput | 100% baseline | 90-95% | 5-10% lower |
| Setup complexity | Lower | Higher | Significant |
| Ecosystem | HPC-focused | Enterprise-friendly | Different skill sets |
| Cost | Higher | 60% of InfiniBand | Major savings |
Based on industry benchmarks and user reports
When to Choose InfiniBand
Choose InfiniBand when:
- You need absolute maximum performance
- Your team has InfiniBand expertise
- You’re running tightly-coupled HPC simulations
- Simplicity matters more than cost
- You want single-vendor support (NVIDIA/Mellanox)
When to Choose RoCEv2
Choose RoCEv2 when:
- Cost optimization is a priority
- You have Ethernet networking expertise
- The 5-10% performance difference is acceptable
- You want commodity switching hardware
- You’re building at scale (costs multiply)
The tuning factor: RoCEv2 performance depends heavily on proper configuration. PFC, ECN, buffer settings, congestion control—these require expertise. Poorly tuned RoCEv2 delivers 70% of InfiniBand. Properly tuned RoCEv2 hits 90-95%.
Reddit insight: “We switched from InfiniBand to RoCEv2 last year. 95% of the performance at half the cost. QSFP-DD modules work for both.”
QSFP-DD Module Selection for AI
Which 800G module for your AI cluster?
Not all 800G modules suit AI workloads. The choice depends on distance, fiber type, and whether you’re mixing speeds during migration.
SR8 for AI Clusters
SR8 covers 100 meters over OM4 multimode fiber. It’s the workhorse for dense GPU rack connections.
Specifications:
- Reach: 100m over OM4
- Connector: MTP-12 APC
- Power: 12-15W
- Best for: Intra-rack GPU-to-switch connections
In AI clusters, SR8 connects GPUs to leaf switches within the same rack. The short reach matches the dense rack designs typical in AI infrastructure—40+ GPUs per rack.
DR8 for AI Spine
DR8 stretches to 500 meters over OS2 single-mode fiber. This is your spine-to-leaf workhorse.
Specifications:
- Reach: 500m over OS2
- Connector: MTP-12 APC
- Power: 15-18W
- Best for: Spine-leaf connections within data center
The 500m reach handles most intra-building spine connections. DR8 uses single-mode optics for better signal integrity at distance.
2FR4 for Migration
2FR4 reaches 2km and naturally breaks out to two 400G connections. This is your migration friend.
Specifications:
- Reach: 2km over OS2
- Connector: LC duplex
- Power: 16-20W
- Best for: Mixed 400G/800G environments
Why 2FR4 matters for AI: When you’re migrating from 400G to 800G, 2FR4 lets you connect 800G spine switches to 400G leaf switches without separate aggregation layers. One 800G port becomes two 400G ports. Clean and simple.
Module Selection Matrix
| Use Case | Recommended Module | Why |
| GPU to leaf (intra-rack) | SR8 | Short reach, lower cost |
| Leaf to spine | DR8 | Medium reach, spine standard |
| Mixed 400G/800G | 2FR4 | Breakout capability |
| Storage network | DR8 or FR4 | Depends on distance |
Need help selecting modules? See our 400G QSFP-DD module types guide for detailed specifications.
Deployment Planning
Your AI cluster deployment checklist.
Before you order hardware, work through this checklist. Each item catches common deployment issues that delay projects.
Pre-Deployment Checklist
GPU Count and Bandwidth Calculation
Count total GPUs in cluster
Multiply by bandwidth per GPU (400G or 800G)
Calculate aggregate bandwidth requirement
Verify spine capacity can handle aggregation
Switch Port Density Planning
Calculate leaf switch count based on GPU density
Calculate spine switch count based on leaf count
Verify 800G port availability on spine switches
Plan for 20% growth headroom
Cable Type and Length Requirements
Measure distances between racks
Choose SR8 for <100m, DR8 for <500m
Account for cable routing (not straight-line distance)
Plan cable management for high fiber counts
Power and Thermal Verification
Calculate total module power consumption
Add switch power consumption
Verify data center cooling capacity
Check rack power distribution limits
For detailed power planning, see our QSFP-DD power and thermal guide.
QSFP-DD Verification
Switch Compatibility Check
Verify switch supports 800G QSFP-DD
Check firmware version for 800G support
Confirm backward compatibility with 400G if needed
Validate buffer sizes for lossless operation
Module Compatibility Matrix
| Switch Platform | 800G SR8 | 800G DR8 | 800G 2FR4 |
| Arista 7060X5 | ✓ | ✓ | ✓ |
| Cisco 8100 | ✓ | ✓ | ✓ |
| Juniper PTX | ✓ | ✓ | Limited |
| NVIDIA SN5600 | ✓ | ✓ | ✓ |
Check specific switch datasheets for latest compatibility
Cost Considerations
TCO for 256-GPU cluster: QSFP-DD vs OSFP.
AI networking costs multiply quickly. A 256-GPU cluster needs roughly:
- 256 NICs (one per GPU)
- 16 leaf switches
- 8 spine switches
- 512+ optical modules
- Miles of fiber cabling
Cost Breakdown by Component
NIC Costs
- ConnectX-7 (400G): ~$1,200 each
- ConnectX-8 (800G): ~$2,400 each
Switch Costs
- 64-port 800G switch: 80,000−80,000−120,000
- Vendor varies pricing significantly
Module Costs
- 800G SR8 QSFP-DD: 600−600−900
- 800G DR8 QSFP-DD: 800−800−1,200
- 800G OSFP modules: Similar pricing
Cabling Costs
- MTP-12 OM4: 50−50−100 per cable
- MTP-12 OS2: 80−80−150 per cable
- Costs scale with fiber count
QSFP-DD Cost Advantages
The backward compatibility of QSFP-DD creates real savings:
- Infrastructure reuse: Connect to existing 400G switches without adapters
- Mixed-speed flexibility: Run 400G and 800G on same switch fabric
- Reduced sparing: One module type covers multiple speeds
- Lower training costs: Team already knows QSFP-DD
Real-world example: A 256-GPU cluster using QSFP-DD saved $180,000 compared to OSFP by:
- Reusing 8 existing 400G switches ($80,000)
- Avoiding adapter modules ($45,000)
- Reduced sparing inventory ($35,000)
- Lower training/consulting costs ($20,000)
For detailed TCO analysis, see our QSFP-DD pricing and TCO guide.
FAQ
How many GPUs before I need 800G?
The minimum number of GPUs required to meet the 800G threshold stands at 128 GPUs. At 128 H100s, you need 51.2T of aggregate bandwidth. The 800G spine connections enable reduced switch requirements and simplified system design. The 400G spine system functions properly until the point where 128 GPUs operate together. The 800G system becomes necessary when 512 GPUs operate in a system.
Which option should I select between QSFP-DD and OSFP for my AI training needs?
Select QSFP-DD when your organization possesses 400G systems or operates multiple types of workloads. Select OSFP for dedicated training clusters which require no backward compatibility with existing systems. QSFP-DD provides users with multiple options. OSFP provides superior thermal efficiency while maintaining complete compatibility with NVIDIA systems.
Which network protocol should I choose between RoCEv2 and InfiniBand for LLM training?
RoCEv2 provides 90-95% of InfiniBand performance at 60% of cost. The best option for your needs is InfiniBand because it delivers maximum performance while making system administration easy. Choose RoCEv2 for cost optimization, especially at scale. Proper RoCEv2 tuning (PFC, ECN, buffers) is essential for good performance.
Can QSFP-DD handle both compute and storage?
Yes, the system requires proper planning to achieve its goals. Use 800G QSFP-DD for the compute backend fabric. Use 400G QSFP-DD for storage networks (which usually meet the required capacity). The two systems must maintain separate operations even though they use the same physical switches. The storage system can handle oversubscription, while the compute system requires dedicated resources.
Which system provides better cost value between QSFP-DD and OSFP?
Both modules have matching pricing. The savings come from the ability to reuse existing infrastructure. The implementation of QSFP-DD enables backward compatibility with 400G, which allows a cost reduction of 150,000−150,000−300,000 for a 256-GPU cluster through the reuse of existing switches without needing adapters.
What circumstances require my transition from 400G to 800G?
You should perform the upgrade once your organization reaches its bandwidth capacity or when it requires development of new clusters that will exceed 128 GPUs. You should choose 800G for your next expansion after your current system reaches 400 GB. The QSFP-DD system enables you to combine 400G and 800G technology during the migration process, which minimizes the risks associated with system upgrades.
Conclusion
Sarah achieved her $270,000 savings through her knowledge of AI networking trade-offs. RoCEv2 on QSFP-DD delivers most of InfiniBand’s performance at a fraction of the cost. The 5-10% performance difference rarely justifies the 40% price premium.
Key takeaways:
- Size for 800G at 128+ GPUs—the math is relentless as clusters scale
- QSFP-DD offers flexibility—backward compatibility saves real money
- RoCEv2 is viable for most AI training—90% of performance at 60% of cost
- Plan three networks—backend, storage, and frontend each have different needs
- Don’t forget storage bandwidth—starved GPUs waste money
The 1.6T horizon is coming. QSFP-DD1600 will fit the same ports as 800G. Your investment in QSFP-DD infrastructure protects forward.
Ready to plan your AI cluster networking? Contact FiberMall’s technical team for cluster sizing, module selection, and deployment guidance. We’ve helped dozens of organizations architect QSFP-DD AI infrastructure—from 64-GPU training clusters to 1,024-GPU AI factories.
Explore our 800G QSFP-DD product line for AI cluster deployments.
The AI infrastructure decisions you make today will determine your training capabilities for years. Choose wisely.
Related Products:
-
QSFP-DD-800G-DR8 800G-DR8 QSFP-DD PAM4 1310nm 500m DOM MTP/MPO-16 SMF Optical Transceiver Module
$1000.00
-
QSFP-DD-800G-FR8 QSFP-DD 8x100G FR PAM4 1310nm 2km DOM MPO-16 SMF Optical Transceiver Module
$1200.00
-
QSFP-DD-800G-2FR4 800G QSFP-DD 2FR4 PAM4 1310nm 2km DOM Dual CS SMF Optical Transceiver Module
$1900.00
-
QSFP-DD-800G-2FR4L QSFP-DD 2x400G FR4 PAM4 CWDM4 2km DOM Dual duplex LC SMF Optical Transceiver Module
$1800.00
-
QSFP-DD-800G-DR8D QSFP-DD 8x100G DR PAM4 1310nm 500m DOM Dual MPO-12 SMF Optical Transceiver Module
$1000.00
-
QSFP-DD-800G-FR8L QSFP-DD 800G FR8 PAM4 CWDM8 2km DOM Duplex LC SMF Optical Transceiver Module
$3000.00
-
QSFP-DD-400G-SR8 400G QSFP-DD SR8 PAM4 850nm 100m MTP/MPO OM3 FEC Optical Transceiver Module
$149.00
-
QSFP-DD-400G-DR4 400G QSFP-DD DR4 PAM4 1310nm 500m MTP/MPO SMF FEC Optical Transceiver Module
$400.00
-
QSFP-DD-400G-SR4 QSFP-DD 400G SR4 PAM4 850nm 100m MTP/MPO-12 OM4 FEC Optical Transceiver Module
$450.00
-
QSFP-DD-400G-FR4 400G QSFP-DD FR4 PAM4 CWDM4 2km LC SMF FEC Optical Transceiver Module
$500.00
-
QSFP-DD-400G-XDR4 400G QSFP-DD XDR4 PAM4 1310nm 2km MTP/MPO-12 SMF FEC Optical Transceiver Module
$550.00
-
QSFP-DD-400G-LR4 400G QSFP-DD LR4 PAM4 CWDM4 10km LC SMF FEC Optical Transceiver Module
$600.00
-
QSFP-DD-400G-SR4.2 400Gb/s QSFP-DD SR4 BiDi PAM4 850nm/910nm 100m/150m OM4/OM5 MMF MPO-12 FEC Optical Transceiver Module
$750.00
-
QSFP-DD-400G-ER4 400G QSFP-DD ER4 PAM4 LWDM4 40km LC SMF without FEC Optical Transceiver Module
$3500.00
-
QSFP-DD-400G-ER8 400G QSFP-DD ER8 PAM4 LWDM8 40km LC SMF FEC Optical Transceiver Module
$3800.00
-
QSFP-DD-400G-LR8 400G QSFP-DD LR8 PAM4 LWDM8 10km LC SMF FEC Optical Transceiver Module
$3800.00
Related Posts
- QSFP-DD Troubleshooting Guide: Fix 400G/800G Link Issues Fast
- Overview of 400G QSFP-DD Optical Transceiver Module
- 400G Data Center: QSFP-DD AOC Cables
- On 400G QSFP-DD SR8 Optical Transceiver Module
- Thermal design study of 200G QSFP-DD LR4 optical module
- 800G QSFP-DD is coming after 400G
- Core Technologies in 400G QSFP-DD AOC: PAM4 and DSP
- 400G QSFP-DD Transceiver in Data Center: Types and Wiring Scheme
- Demand for 400G QSFP-DD Optical Transceiver to Rise in 2021
