Blog

A place to share and bookmark notes on GPU scheduling, capacity planning for AI workloads, and the infrastructure behind large-scale ML.

All opinions are my own and do not represent those of my employer.

Part 0 — Foundations: The Utilization Problem and a Map of the Stack
Part 1 — Anatomy of a Compute Unit: What One GPU Server Gives You
Part 2 — Sharing Within One GPU: Small Jobs Like Inference
Part 3 — Placing Jobs Across the Cluster: Multi-GPU Training
Part 4 — Sharing the Cluster Across Teams and Time: Reclaiming Idle Capacity
Part 5 — Failure Recovery: Keeping Placed Work Alive
Part 6 — Measurement and Economics: What Utilization Is and What Idle Costs
Part 7 — Synthesis