The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Baseten's team walks a 200,000-token request through a production inference stack: cache lookup first, then disaggregated prefill and decode across separate GPU pools, with a speculative decoder trained on coding traffic sitting in front. The best single explanation in the catalog of what happens after you call the API.