Reference architectures, explained

The workload changes. So does the right architecture.

Explore how servers, trays, rack scale-up domains and external fabrics relate. Hardware references establish boundaries; the interactive model makes its derivations and unknowns visible. It is not a benchmark, certification or procurement-ready bill of materials.

Interactive architecture example

Follow the work.
Understand the fabric.

Published reference counts, explicit modeling assumptions. No live equipment, procurement certification or benchmark.

Placement, failure & demand assumptions

NVIDIA DGX GB300 NVL72 · InfiniBand XDR

NVIDIA DGX GB300 NVL7272 GPUs per rack scale-up domainInfiniBand XDR · separate client, storage and management networks

Identity & connection inspector

Select a labeled component or bundle. Every bundle retains its relationship IDs.

Reference, assumptions & limitations

NVIDIA DGX GB300 NVL72 — schematic shared scale-up domain; not a physical switch-port map

Transfer time ≈ fixed path overhead + bytes / effective bandwidth + queueing. Raw link capacity is not workload throughput. Hop counts depend on the chosen route.

Start with the scale-up boundary

An eight-GPU DGX server is not interchangeable with a 72-GPU rack. CPUs, trays, accelerator packages and workload ranks are separate concepts. Selecting a reference rebuilds the model; selecting a workload changes traffic over installed wiring.

Training and inference share the physical fabric

DP synchronizes replicas; sharding distributes and gathers ownership; TP stays in the selected scale-up domain; pipeline stages exchange activations and gradients; MoE dispatches and combines expert work. These are teaching relationships, not an exact collective-library schedule.

Follow the client request and returned tokens

Aggregated inference keeps prefill and decode together, within one domain or a distributed replica. Disaggregated inference separates prefill and decode pools; KV transfer can stay local or cross scale-up domains. The frontend service network is separate from the backend compute fabric.

Latency, bandwidth and high availability

Transfer time is approximately fixed path overhead plus bytes divided by effective bandwidth plus queueing. Smaller transfers can be sensitive to latency; larger collectives and KV transfers can be sensitive to bandwidth. A surviving route does not establish session continuity, measured failover time or workload throughput.

The source is part of the design

References are pinned by retrieval date and content hash. Published, preliminary, derived and unknown fields are distinguished. AMD MI3XX options use Pollara 400 with standard RoCEv2; Helios uses Vulcano 800 with a selected standard-RoCEv2 logical design and a separate UALoE scale-up domain. Exact OEM board population and physical pairing still require installation evidence. Physical cages, cable assemblies, length, rack positions and prices are not inferred from diagram geometry.

Keep the original example in context

The earlier 4,096-GPU H100-style model is an illustrative three-tier graph. Its URL and stored design revisions retain that meaning; it is not relabeled as a rack reference.

A useful next step

Review a platform decision

A rough outline is enough to start a scope conversation.

Review a platform decision ↗

Content revised 2026-09-29.