What Is Tail Latency? p99 and Slow Requests, Explained
Tail latency is the response time of your slowest requests, measured as p99 or p999. Why averages hide it, what causes it, and how to reduce it.
Tail latency is the response time of the slowest fraction of requests a system serves, usually reported as a high percentile such as p99 (the time under which 99% of requests complete) or p999 (99.9%). Averages and medians describe the typical request; tail latency describes the bad ones. In large distributed systems, the bad ones are not rare edge cases. They are what a meaningful share of users actually experience, and they often determine how fast the whole system feels.
Percentiles, briefly
If you sort every request’s latency over a time window, the pth percentile is the value at position p% through the list.
- p50 (median): half of requests are faster, half slower. The “typical” experience.
- p90 / p95: the start of the tail.
- p99: one request in a hundred is slower than this.
- p999: one in a thousand.
The mean is a poor summary for latency because latency distributions are skewed. Most requests cluster near a fast floor, while a long tail stretches to the right. A handful of multi-second requests barely moves the mean but is exactly what users complain about. That is why service-level objectives are usually written in percentiles: “p99 under 300 ms” rather than “average under 100 ms.”
Why the tail matters more than it seems
Fan-out amplifies it
Modern backends rarely answer a request with a single call. A page might query a dozen services, and a search query might be scattered to hundreds of index shards. The response is only ready when the slowest dependency replies.
The arithmetic is unforgiving. If each backend call independently has a 1% chance of being slow, a request that waits on 100 such calls has about a 63% chance (1 − 0.99¹⁰⁰) of hitting at least one slow call. A backend’s p99 becomes the front end’s median. This is why teams operating at scale obsess over the tail: in a fan-out architecture, the tail is the user experience.
Heavy users hit it more
The users who make the most requests per session, and are often the most valuable, are statistically the most likely to experience the tail on any given visit.
What causes tail latency?
Slow outliers rarely come from the code path itself. They usually come from contention and interference:
- Queueing. When requests arrive faster than a server can process them, even briefly, they wait. Queueing delay grows sharply as utilisation approaches 100%, so a server running “hot” has a much fatter tail than one with headroom.
- Garbage collection pauses. Managed runtimes occasionally stop or slow application threads to reclaim memory. Requests in flight during a pause absorb the whole delay. See JavaScript garbage collection explained for how this works in one runtime.
- Cache misses. A request that misses a cache and goes to disk or a database is far slower than one that hits.
- Noisy neighbours. Other workloads on the same host compete for CPU, memory bandwidth, disk, and network.
- Background work. Compaction, log rotation, backups, and periodic jobs steal resources at unpredictable moments.
- Network effects. Packet loss triggers TCP retransmission timeouts, which can add far more delay than the original round trip.
- Cold starts. New containers, JIT warm-up, and freshly opened connections are slower than warm ones.
- Lock contention. Requests that need the same row or mutex serialise behind each other.
How to measure it correctly
Getting tail numbers right is surprisingly easy to get wrong.
- Never average percentiles. The mean of ten servers’ p99s is not the fleet’s p99. Aggregate the raw data or use mergeable histograms, then compute percentiles.
- Use histograms, not just summaries. Bucketed latency histograms can be combined across hosts and time windows; precomputed percentiles cannot.
- Watch for coordinated omission. A load generator that waits for each response before sending the next request quietly stops sending during a stall, so the stall is under-sampled. Good benchmarking tools issue requests on a fixed schedule and measure from the intended send time.
- Measure where users are. Server-side timing misses network and queueing delays in front of the server. Pair it with client-side measurement, as discussed in our guide to observability.
Techniques to reduce tail latency
Keep utilisation below the knee
Because queueing delay rises steeply near saturation, leaving capacity headroom is the simplest tail-latency tool. Autoscaling on p99 latency or queue depth, not just CPU, helps.
Hedged requests
Send a request to one replica, and if it has not answered within, say, the p95 latency, send a second copy to another replica. Use whichever answer arrives first and cancel the other. Because the hedge fires only for slow requests, it adds only a few percent of extra load while sharply cutting the tail. A variant, the tied request, sends to two replicas at once and lets them cancel each other when one starts work.
Timeouts, retries, and backoff
Bounded timeouts prevent one stuck dependency from holding a request forever. Retries recover from transient slowness, but must use exponential backoff with jitter and a retry budget, or they become a load amplifier during incidents.
Fail fast and shed load
A circuit breaker stops calling a dependency that is clearly unhealthy, and load shedding rejects excess work early rather than letting every request queue. A quick error is often better for users than a slow success.
Smarter load balancing
Round-robin ignores how busy a server is. Strategies such as least-outstanding-requests or “power of two choices” (pick two servers at random, send to the less loaded one) route around slow instances. Our load balancer explainer covers these algorithms.
Reduce variance at the source
Tune GC settings, pre-warm caches and connection pools, isolate background jobs onto separate resources, and pin latency-sensitive workloads away from batch work.
The takeaway
Tail latency is the experience of your slowest requests, measured at p99 and beyond. Averages hide it, fan-out architectures amplify it, and it usually comes from queueing and interference rather than slow code. Measure it with mergeable histograms and honest load tests, keep capacity headroom, and use hedged requests, bounded retries, and load-aware balancing to keep the tail short.
Keep reading
Chisato · · 4 min read What Are Vector Clocks? Ordering Distributed Events
A vector clock is a per-node counter array that lets distributed systems tell whether one event happened before another, without a shared clock.
The Lycoris Team · · 4 min read Quorum Consensus Explained: N, W, and R
Quorum consensus lets distributed databases tune consistency and availability by requiring reads and writes to touch overlapping subsets of replicas.
Chisato · · 4 min read What Is Replication Lag? Causes and Fixes
Replication lag is the delay between a write on the primary database and its arrival on a replica. What causes it, how to measure it, and how to reduce it.