Start with the scale-up boundary
An eight-GPU DGX server is not interchangeable with a 72-GPU rack. CPUs, trays, accelerator packages and workload ranks are separate concepts. Selecting a reference rebuilds the model; selecting a workload changes traffic over installed wiring.
Training and inference share the physical fabric
DP synchronizes replicas; sharding distributes and gathers ownership; TP stays in the selected scale-up domain; pipeline stages exchange activations and gradients; MoE dispatches and combines expert work. These are teaching relationships, not an exact collective-library schedule.
Follow the client request and returned tokens
Aggregated inference keeps prefill and decode together, within one domain or a distributed replica. Disaggregated inference separates prefill and decode pools; KV transfer can stay local or cross scale-up domains. The frontend service network is separate from the backend compute fabric.
Latency, bandwidth and high availability
Transfer time is approximately fixed path overhead plus bytes divided by effective bandwidth plus queueing. Smaller transfers can be sensitive to latency; larger collectives and KV transfers can be sensitive to bandwidth. A surviving route does not establish session continuity, measured failover time or workload throughput.
The source is part of the design
References are pinned by retrieval date and content hash. Published, preliminary, derived and unknown fields are distinguished. AMD MI3XX options use Pollara 400 with standard RoCEv2; Helios uses Vulcano 800 with a selected standard-RoCEv2 logical design and a separate UALoE scale-up domain. Exact OEM board population and physical pairing still require installation evidence. Physical cages, cable assemblies, length, rack positions and prices are not inferred from diagram geometry.
Keep the original example in context
The earlier 4,096-GPU H100-style model is an illustrative three-tier graph. Its URL and stored design revisions retain that meaning; it is not relabeled as a rack reference.