4K means GPUs, not servers
This illustrative compute-only graph contains 512 servers with eight GPUs and eight GPU-affine 400 Gb/s NIC ports each. Sixteen blocks contain 32 servers each. It is structurally checked offline, not hardware-tested, vendor-certified or a production bill of materials.
- 128 rail-aligned leaves: each uses 32 host-facing and 32 spine-facing logical 400G ports.
- 16 spine groups (two banks × eight rails), with eight spines per group: 128 spines. Each leaf has four links to each of its eight spines.
- 64 cores in eight groups. Spine slot k connects to eight cores in group k with four links per pair. Each core uses 64 ports, receiving four links from one spine in each of 16 groups.
- 320 compute switches; 4,096 endpoint–leaf + 4,096 leaf–spine + 4,096 spine–core = 12,288 logical links.
- All compute ports are consumed. Real designs need fabric management, spares, growth, storage/frontend/OOB networks, supported media and routing. Counts can change.
Logical ports are not cable assemblies
The example uses a QM9700-class radix: 64 logical 400G ports exposed through 32 twin-port OSFP cages. Logical numbering is abstract, not vendor interface syntax. Cages, fiber strands, logical links and purchased harnesses have different counts.
- Per-server raw endpoint injection: 8 × 400 Gb/s = 3.2 Tb/s, one direction.
- Each leaf: 12.8 Tb/s host-side and 12.8 Tb/s uplink-side, raw.
- All endpoints: 1,638.4 Tb/s = 1.6384 Pb/s raw injection, one direction. Half: 819.2 Tb/s, a planning cut quantity, not a blanket bisection guarantee.
- Do not add opposite directions as useful throughput. NCCL algorithm bandwidth, normalized bus bandwidth and useful application throughput are different measures.
The fabric should match the communication pattern—not just the GPU count.
An illustrative TP=8 × PP=8 × DP=64 mapping covers 4,096 ranks. TP stays in the eight-GPU server scale-up domain here; pipeline activations cross nodes; same-shard gradients can use rail-aligned paths. Shared cores connect groups and rails: these are not eight disconnected physical networks.
- Replicated data parallelism: all-reduce. Sharded optimizer: reduce-scatter and all-gather are separate stages, not the same algorithm.
- MoE dispatch/combine is all-to-all between selected experts. Expert parallelism is not a fourth independent multiplier in TP × PP × DP.
- Context parallelism and long sequences add communication. Collective size/concurrency, overlap, GPU/NIC/NUMA affinity and placement matter.
- Checkpoint/data-feed interference, congestion and stragglers can expose communication. Animation is a selected logical path schematic, not a packet-accurate NCCL schedule.
Training coordinates work. Inference also meets a request latency budget.
Aggregated serving routes requests to independent replicas or replica groups. Prefill and decode may share workers; the entire fleet need not synchronize on every request. Disaggregated serving separates prefill and decode pools, transferring KV cache to selected decode workers before streaming output.
- Prefill is often compute-intensive; decode may be memory-bandwidth or capacity-sensitive. Model, batch, concurrency, context and runtime determine behavior.
- KV transfer and pool fragmentation can outweigh disaggregation benefits. Aggregated serving can be simpler or better. Large TP and MoE inference can still need substantial low-latency bandwidth.
- Time to first token: request-to-first-output delay. Inter-token latency: delay between output tokens. Time per output token: average per-token generation interval.
- p95/p99 describe latency percentiles. Throughput within an SLO counts work meeting the stated objective. KV capacity bounds resident cache. No measured values or universal prefill/decode ratio are asserted.
Failure paths and separate operational layers
Select the illustrative spine failure to show a remaining path through another slot. Graph reachability survives one spine removal, with reduced capacity and changed path choice. Routing convergence and failover performance have not been measured.
- A single host NIC-to-leaf link is not made redundant by having extra spines. A lost rail and shared failure domains need separate planning.
- Scale-up is distinct from compute InfiniBand. Storage/checkpoint, frontend/service Ethernet and out-of-band management overlays are conceptual and not sized in the port budget.
- Validate cable type, link width/rate, firmware/software compatibility and management connectivity. InfiniBand credit/fabric/congestion management is not Ethernet PFC/ECN/RoCE configuration.
- The view samples named paths from the graph; hiding a layer does not delete connectivity. No routing, latency, collective performance or customer result is certified.