The 99th Percentile Isn't a Number: It's a Promise
There's a moment in every on-call rotation where you realize the p50 looks perfect, the p95 is fine, but the p99 is a cliff. Requests that should take 50 milliseconds are taking 4 seconds. Users are seeing spinner hell. Your p99 SLA is blown, and nobody can tell you why. This is the story of where latency budgets actually break—and what to do about it. It's not about tweaking a few configs. It's about making a decision, and making it before the pager goes off. Let's get honest about what's in your control, what isn't, and what to do when the tail winds start blowing. Who Owns the Tail? The Decision You Can't Defer Why SRE and dev teams often fight over p99 ownership The argument usually starts in a post-incident review. Someone pulls up the latency graph, points at the spike, and says “this is a networking problem.