High-Performance AI Systems
Engineering Hyperscale GPU Clusters
Most material on large-scale AI either stops at the model or starts at the datacentre door. This book covers the space in between: what a training job does to the hardware it runs on, how a cluster has to be built to absorb it, and what breaks first when you scale it up.
It is written for ML engineers who have outgrown a single node and for infrastructure engineers who inherited a GPU fleet and need to know which knobs matter.
Foundations
What AI workloads actually demand of hardware - how training and inference consume compute, memory, and bandwidth, and where the ceilings come from.
Infrastructure
The cluster itself: GPU architecture, scale-up and scale-out interconnect, storage, power, and the physical layer that ties them together.
Operations
Running it in production - distributed training at scale, failure modes, observability, and the operational practice that keeps utilisation high.