How does InfiniBand topology scale?

How does InfiniBand topology scale?

In 2026, the scale of AI "super-factories" has pushed InfiniBand topology from simple racks to massive, multi-thousand-node clusters. While Ethernet is excellent for general cloud use, InfiniBand remains the gold standard for scaling because its topologies are designed to be lossless and predictable at extreme densities.

Here is how InfiniBand scales through its primary architectural blueprints.


1. The Fat Tree (Folded Clos)

The Fat Tree is the most common topology in OCI and NVIDIA SuperPODs. It is a hierarchical design where the "branches" (bandwidth) get thicker as you move toward the core.

  • Non-Blocking Scaling: A "Full Bisection Bandwidth" Fat Tree ensures that every node can talk to every other node at full speed (e.g., 400Gbps NDR) simultaneously.

  • The 3-Tier Limit: Most enterprise clusters use a 3-tier Fat Tree (Leaf → Spine → Core).

    • Scale: Using modern 64-port switches (like the Quantum-2), a 3-tier Fat Tree can scale to over 27,000 GPUs in a non-blocking configuration.

  • The Oversubscription Lever: To save money, architects can "thin" the upper tiers (e.g., a 2:1 blocking factor). This allows for even more nodes, but at the cost of potential congestion when all nodes try to communicate at once.


2. Dragonfly+ (The "Direct Path" Scale)

As clusters push beyond 30,000 nodes, the sheer number of cables in a Fat Tree becomes a physical nightmare. Dragonfly+ is the 2026 solution for ultra-large scale-out.

  • Grouped Architecture: Routers are divided into "groups." Within a group, everything is tightly connected. Between groups, there are "global links."

  • Extreme Scalability: A Dragonfly+ topology can connect over 1 million nodes in just 3 hops.

  • Cost Efficiency: It uses significantly fewer cables and switches than a Fat Tree for the same number of nodes.

  • The Catch: It relies heavily on Adaptive Routing. Since there are fewer global paths, the network must be smart enough to reroute traffic instantly if one "global" link gets jammed.


3. Rail-Optimized Scaling

For AI training, we don't just scale nodes; we scale GPU rails.

In an OCI Bare Metal instance with 8 GPUs, each GPU has its own dedicated InfiniBand NIC. To optimize performance, we build "Rails":

  • Rail-Alignment: NIC 0 from every server is connected to Switch 0. NIC 1 is connected to Switch 1.

  • The Scaling Secret: When GPUs perform an All-Reduce operation (syncing their "learning"), they only talk to their counterparts on the same "rail." This reduces "cross-talk" and ensures that scaling from 128 to 1,000 nodes doesn't result in a performance drop.


4. In-Network Computing (SHARPv3)

In 2026, InfiniBand doesn't just "move" data; it "scales" the math.

  • SHARP (Scalable Hierarchical Aggregation and Reduction Protocol): In traditional networks, if 1,000 nodes want to sum their data, they send it all to one node (a bottleneck).

  • The Logic Move: With SHARPv3, the switches themselves perform the math. As data passes through the switch fabric, the switches sum the numbers and pass only the result up the tree.

  • Result: This reduces the amount of data traveling through the top of the topology by orders of magnitude, allowing the cluster to scale without saturating the core links.


Summary: Scaling Comparison

TopologyBest ForMax Scale (2026)Complexity
2-Tier Fat TreeSmall/Mid Clusters~2,000 NodesLow
3-Tier Fat TreeEnterprise AI (SuperPOD)~27,000+ NodesHigh (Cabling)
Dragonfly+Exascale / National Labs1,000,000+ NodesModerate (Logic)
Torus (3D/6D)Specific HPC PhysicsUnlimitedVery High

Key Takeaway for Your Blog:

"Scaling InfiniBand is no longer just about buying more switches; it's about choosing the right geometry for your data. For most AI workloads, a Rail-Optimized Fat Tree provides the predictable performance needed for training, while Dragonfly+ offers the radical scale required for the next generation of global AI foundations."

Looking for servers Rental ?

Call Our Expert :


  • (call for rental enquiries)

Email us :