Backend & systems·Scaling and load management
the median request is 40ms, the 99th percentile is four seconds, and every page that fans out to ten services hits it.
Tail latency
Also calledp99, long tail, tail at scale
The slow end of the latency distribution, where garbage collection, cache misses, retries and contention live. It matters disproportionately because a request that fans out to N services experiences the tail with probability roughly N times higher. Averages hide it completely, which is why percentiles are the only useful summary.