Orientation
An AI cluster is a data-movement machine
Installed compute becomes productive only when every dependency reaches it on time.
A model does not run on arithmetic alone. Training samples must stream from storage, parameters must be available in memory, partial results must cross links, and updates must be synchronized before the next step can begin. The fabric is the complete delivery system for that motion: on-board buses, scale-up links, rack networks, campus fiber, storage paths, drivers, collective libraries, schedulers, and the operational controls that hold them together.
This changes how the buildout should be read. A powerful accelerator waiting for a packet is an expensive idle asset. A fast switch attached to a poorly shaped workload is equally wasted. Effective capacity therefore emerges from the slowest recurring boundary between compute, memory, network, and storage—not from the largest component specification printed on a product sheet.
Reader's rule: Follow the byte, not the brand. Ask where data starts, which boundaries it crosses, how often it moves, and what happens when one path fails.
An inference request creates several traffic classes
- Admit and route— Authenticate the request, apply tenant policy, select a model replica, and preserve an end-to-end deadline.
- Load and prefill— Make weights and adapters resident, process prompt tokens, and create key-value state near the chosen workers.
- Transfer and decode— Move or reuse state, schedule token generation, and prevent queueing from defeating interactive latency.
- Stream and observe— Return tokens while recording latency, cache use, errors, retries, energy, and tenant attribution.



