nikhilshetty.net

High-Performance AI Systems

Engineering Hyperscale GPU Clusters

Nikhil Shetty and Rik Kisnah  |  Packt  |  13 chapters, in progress

Most material on large-scale AI either stops at the model or starts at the datacentre door. This book covers the space in between: what a training job does to the hardware it runs on, how a cluster has to be built to absorb it, and what breaks first when you scale it up.

It is written for ML engineers who have outgrown a single node and for infrastructure engineers who inherited a GPU fleet and need to know which knobs matter.

  • Foundations

    What AI workloads actually demand of hardware - how training and inference consume compute, memory, and bandwidth, and where the ceilings come from.

  • Infrastructure

    The cluster itself: GPU architecture, scale-up and scale-out interconnect, storage, power, and the physical layer that ties them together.

  • Operations

    Running it in production - distributed training at scale, failure modes, observability, and the operational practice that keeps utilisation high.