I read Distributed AI Systems by Fuheng Wu to better assess the infrastructure behind AI applications. My work is closer to agent systems than training or inference stacks; I wanted a clearer understanding of the constraints that shape their performance.
Disclosure: Packt provided a review copy of this book.
Memory requirements were one of those questions. Calculating the size of a model’s weights is straightforward enough: you need the parameter count and the number format used to store them. The less obvious requirement is everything that must fit alongside those weights, especially when a language model processes a long conversation.
The book introduces that problem early and returns to it in considerably more detail in the inference chapters. During generation, the model stores keys and values computed for earlier tokens in what is called the KV cache. It can then reuse those calculations as it produces the next token. The cache grows as the sequence gets longer, so the prompt and the generated answer both contribute to its memory requirements. Several concurrent requests need more memory again.
The practical consequence is that context length and KV-cache memory are directly related. A model’s advertised context window tells you the maximum length it supports; the cache required for a particular request depends on how much of that context is actually used, as well as the model architecture and serving implementation. A model may support a long conversation without your hardware having enough memory to run it.
I liked that the book gave me a way to reason through this. Knowing that long contexts are expensive is easy. Understanding where the expense comes from makes the fact much more useful.
Once you start accounting for that memory, the hardware discussion becomes easier to appreciate. Capacity is only part of the problem. The system also has to move data quickly enough to keep its processing units occupied, and moving data has costs of its own. The explanation of systolic arrays helped here. These are grids of processing elements that pass data between neighboring elements and reuse it as a calculation progresses. For operations such as matrix multiplication, this reduces repeated memory access and allows many elements to work in parallel.
The architectural explanation also clarifies the reasons for specialized AI hardware, including neural processing units and Google’s TPUs. The book gives you enough detail to understand what these designs are trying to make efficient.
With several devices, you have another question to answer: how should they divide the work? You can give copies of a model different examples, put consecutive layers on different devices, or split calculations within a layer. Mixture-of-experts models offer another division, with different experts placed on different devices. Each arrangement changes what must be stored locally and what must travel between devices.
This was where the hardware and parallelism material came together for me. More GPUs can make a workload possible, but their usefulness depends on how well the work can be divided and how much communication that division requires. The connections between devices become part of the calculation.
The chapter on PyTorch’s Distributed Data Parallel makes the coordination problem concrete. Each GPU works on different examples using its own copy of the model, and the copies synchronize gradients so they continue learning together. From there, the practical details have a clear purpose. Gradient accumulation lets training process a larger effective batch in smaller pieces. Mixed precision reduces memory requirements. Grouping gradients into communication buckets helps overlap their transfer with computation, reducing the time devices spend waiting.
The chapter connects each technique to the constraint it addresses, making the costs of distributed training easier to reason about.
The inference material held my attention most, partly because the memory questions from the beginning become very concrete there. A server has to accommodate requests of different lengths, arriving and finishing at different times. Their caches grow while answers are generated. Memory freed by one request needs to become available to another.
The explanation of vLLM’s PagedAttention makes this problem easy to follow. It divides the KV cache into blocks that can be allocated as needed, so a request does not require one large, continuous region of memory. Continuous batching allows new requests to join while others finish. You can see why these mechanisms affect how many requests a GPU can handle.
SGLang’s RadixAttention adds a possibility that is particularly relevant to assistants: requests can share work when they begin with the same text. Repeated system instructions, examples, or earlier conversation turns may already have cached calculations that a later request can reuse.
That makes request routing more interesting too. A worker with the relevant cached prefix may be able to serve a request more efficiently than one that has to calculate it again. Keeping successive conversation turns on the same worker can therefore have a practical benefit. The book also acknowledges overlapping capabilities between SGLang and vLLM, which I appreciated; the useful question is how these approaches suit a workload.
By this point, serving performance is easier to understand in terms of what an application actually asks the infrastructure to do. Processing a prompt and generating an answer place different demands on the hardware. A request may wait a long time for its first token and then stream quickly, or start promptly and continue slowly. A throughput figure alone cannot describe either experience.
The production-serving and benchmarking chapters follow those distinctions through to routing, reliability, and measurement. Prompt lengths, answer lengths, concurrency, and cache conditions all influence the result. If you want to judge whether a serving configuration works well for your application, your benchmark has to resemble the application.
That is the book’s main value for my work. It has given me a better basis for judging offerings around inference and serving. I can follow their technical claims more closely and assess how they might fit a particular workload. That is useful even when someone else is responsible for running the infrastructure.
For readers working directly in this space, I think there is considerably more to gain. The practical code examples offer a way to learn the material hands-on, across training, fine-tuning, inference, and serving. I particularly appreciated that the author shows how to follow along with multi-GPU examples using freely available infrastructure, making it possible to experiment without buying a machine.
I would recommend it to both groups. If you want to go deep into the stack, there is plenty to work through and put into practice. If, like me, you want a stronger understanding, you can skip some code listings or sections that go further into implementation than you need and still get a great deal from the book.

