Performance Systems Reading Map
This book is the curriculum and experiment notebook. It should not pretend to replace processor manuals, kernel documentation, language standards, or the people who developed the tools being studied.
Use four kinds of material differently:
| Kind | Use it for | Do not assume |
|---|---|---|
| Foundation book | A coherent mental model and vocabulary | Every API detail is current |
| Official documentation | Current contracts, constraints, and configuration | It is a teaching sequence |
| Laboratory repository | Code to run, modify, break, and measure | Its result transfers to your machine |
| Measurement tool | Evidence about a stated experiment | The tool chooses the right question |
Pin the version or commit used in an experiment. Record the CPU, kernel, compiler, flags, NIC, driver, firmware, topology, power policy, and relevant operating-system configuration. “Faster” without that context is not a reusable finding.
If you are buying books
The strongest first purchases for this curriculum are:
- Systems Performance, second edition, for a disciplined whole-system method.
- Understanding Software Dynamics, for explaining intermittent delay, queues, waiting, and long-tail latency with low-overhead tracing.
- The Art of Writing Efficient Programs, for hardware-aware measurement and optimization through C++ experiments.
Then buy according to the layer you are actively studying:
- C++ concurrency: C++ Concurrency in Action, second edition.
- Linux APIs: The Linux Programming Interface.
- Production observability: BPF Performance Tools.
- Deep processor architecture: Computer Architecture: A Quantitative Approach, seventh edition.
Two books are unusually direct about low-latency trading systems:
Use those as architecture tours and implementation prompts, not as the final authority for processor behavior, Linux APIs, venue semantics, or benchmark claims. They connect the layers conveniently; the specialist books and current official documentation establish the details.
A deliberate reading order
1. Learn the modern CPU cost model
Start with Denis Bakhvalov’s open Performance Analysis and Tuning on Modern CPUs. Pair each concept with an exercise from Perf-Ninja: caches, branches, vectorization, memory-level parallelism, and hardware counters become useful only after you predict and measure them.
Agner Fog’s Optimizing software in C++ is a dense reference for compiler behavior and low-level C++ optimization. Treat its advice as hypotheses to verify on the compilers and processors named in your own report.
2. Learn whole-system performance
Brendan Gregg’s Systems Performance, second edition provides the broader method: workloads, CPUs, memory, filesystems, networking, profiling, tracing, and latency outliers. Its most durable lesson is to form a system model before reaching for a favorite tool.
Use BPF Performance Tools and its companion repository when the question requires kernel visibility rather than another application timer.
3. Learn Linux interfaces before bypassing them
Michael Kerrisk’s The Linux Programming Interface is the comprehensive reference. The newer open Linux System Programming Essentials is a shorter practical route through file descriptors, processes, memory, signals, and I/O.
You should be able to explain the normal syscall, scheduler, socket-buffer, and network-stack path before deciding which part to avoid.
4. Learn C++ concurrency as a correctness model
Anthony Williams’s C++ Concurrency in Action, second edition is a useful structured treatment of threads, futures, atomics, the memory model, and lock-free structures. Pair it with current compiler and standard-library documentation: the book targets C++17, while the implementation track here also uses later C++ features.
The objective is not memorizing memory-order names. It is being able to state ownership, invariants, publication edges, and reclamation rules before optimizing a concurrent structure.
5. Move down the network path one layer at a time
“Kernel bypass” is not one feature, and these mechanisms are not substitutes in every workload:
ordinary sockets
↓ batching, affinity, busy polling, steering
io_uring: asynchronous kernel I/O through shared submission/completion rings
↓
XDP / AF_XDP: early packet processing plus shared user/kernel packet rings
↓
DPDK poll-mode drivers: userspace polling and direct NIC descriptor management
- The liburing repository supplies the
reference userspace library and examples for
io_uring.io_uringreduces submission and completion overhead; it is not general kernel bypass. - The Linux kernel’s AF_XDP documentation defines its rings, UMEM ownership, copy modes, and socket behavior.
- XDP Tutorial and BPF examples are laboratories. Follow their own version caveats and use kernel documentation as the current contract.
- The DPDK Programmer’s Guide and poll-mode driver documentation explain a substantially different ownership and operational model.
- Seastar is a valuable C++ laboratory for shared-nothing, one-thread-per-core design. Its tutorial makes futures, sharding, and reactor-style execution concrete.
Do not start with DPDK because it sounds fastest. First measure ordinary sockets, batching, queue placement, and scheduler effects. Every lower-level path trades generality and operational simplicity for more explicit ownership and control.
6. Make the measurements honest
Google Benchmark is a useful C++ harness, not a substitute for experimental design. Use HdrHistogram_c for wide-range latency recording and study its coordinated-omission correction before trusting load-test percentiles.
Final evidence should include distributions, not only averages; warm and cold conditions when both matter; throughput at the reported latency; compiler and binary inspection; and explicit treatment of queueing and coordinated omission.
How the sources map to this book
| Book section | Primary companions | First laboratory |
|---|---|---|
| The Machine | Bakhvalov, Perf-Ninja, Agner Fog | Predict cache and branch behavior, then check counters |
| Operating Systems | Gregg, Kerrisk, BPF tools | Attribute one latency outlier across user and kernel time |
| Concurrency | Williams, Seastar | Prove and measure a bounded SPSC channel |
| Networking and I/O | Kernel AF_XDP docs, liburing, DPDK docs | Trace buffer ownership through four receive paths |
| Performance Engineering | Gregg, Google Benchmark, HdrHistogram | Produce a reproducible latency distribution under load |
| High-Performance C++ | Agner Fog, Williams, compiler output | Compare layout, allocation, and generated code—not language slogans |
The repositories are laboratories, not a giant checklist. A good study cycle is:
- Predict the result from a model.
- Write the smallest experiment that could disprove the prediction.
- Record enough environment detail to reproduce it.
- Inspect counters, traces, or generated code when wall time cannot explain why.
- Change one important variable and repeat.
- Write what would make the conclusion stop being true.