Token Bucket vs Leaky Bucket Rate Limiting
Token bucket allows bursts up to a cap; leaky bucket smooths traffic to a constant rate. How each rate-limiting algorithm works and when to pick it.
Token bucket and leaky bucket are the two classic algorithms behind most rate limiting implementations, and they answer the same question differently: should traffic be allowed to burst, or should it be smoothed to a constant rate? Both are simple enough to implement in a few lines, and the choice between them shapes how your API behaves under load.
Token bucket: allow bursts, cap the average
Picture a bucket that holds tokens, refilled at a fixed rate up to some maximum capacity. Every incoming request consumes one token. If the bucket has tokens available, the request proceeds immediately; if it’s empty, the request is delayed or rejected.
- Bursts are allowed. If the bucket has been filling while idle, a client can spend the whole balance in a single burst — up to the bucket’s capacity — before being throttled.
- Average rate is capped. Over time, the refill rate bounds how many requests get through, since tokens can’t be spent faster than they’re replenished once the bucket is empty.
- Two parameters: bucket size (burst allowance) and refill rate (sustained rate).
This is the algorithm behind most public API rate limits, including the common “1000 requests per hour, but you can fire 50 at once” pattern. It’s forgiving of bursty, realistic client behavior — a user loading a dashboard that fires a dozen requests at once isn’t punished the way a strict per-second cap would punish them.
Leaky bucket: smooth everything to a constant rate
Leaky bucket flips the metaphor: requests pour into a bucket (a queue) and leak out at a fixed, constant rate, regardless of how bursty the input was. If the bucket fills up before it can drain, new requests are dropped or rejected.
- No bursts survive. Even if ten requests arrive simultaneously, they’re processed one at a time at the leak rate — the burst is absorbed into the queue, not passed through.
- Output is always smooth. This is the appeal: whatever’s downstream (a database, a third-party API with its own strict limits) sees a steady, predictable rate of traffic, never a spike.
- Queueing adds latency. Requests that arrive during a full bucket either wait or get dropped; either way, the smoothing comes at the cost of some requests not completing at the moment they arrived.
Leaky bucket is common in network traffic shaping and anywhere a downstream system genuinely cannot tolerate bursts — protecting a fixed-capacity resource, throttling outbound calls to a third-party API with a hard per-second cap, or shaping traffic before it hits a load balancer.
Side-by-side
| Token bucket | Leaky bucket | |
|---|---|---|
| Bursts | Allowed, up to bucket size | Smoothed out entirely |
| Output rate | Variable, capped on average | Constant |
| Implementation | Counter + timestamp | Queue + fixed drain rate |
| Best for | Public APIs, user-facing limits | Traffic shaping, protecting fragile downstreams |
| Rejected requests | Only when bucket is empty | When queue is full |
Where each one actually gets used
Token bucket dominates API gateway and client-facing rate limiting because it matches how real clients behave — nobody sends requests at a perfectly uniform rate, and punishing normal burstiness creates a worse developer experience for no real benefit. It’s also cheap to implement per-client with just a counter and a timestamp, which is why it scales well across a distributed API gateway fleet with each node tracking its own approximate state.
Leaky bucket shows up more in infrastructure and networking contexts — shaping outbound traffic, protecting a downstream service that has a hard ceiling on throughput, or implementing quality-of-service policies on a network link. The queueing behavior is a feature there: you’d rather delay a request by a few hundred milliseconds than let it slam into a fragile backend.
A related pattern worth knowing is the circuit breaker, which handles a different failure mode — not “too much traffic” but “the downstream is already failing” — and the two are often layered together: rate limit to prevent overload, circuit-break to stop calling something that’s already down.
Sliding window as a middle ground
Neither algorithm is the only option. Sliding window counters — tracking request counts over a moving time window rather than a fixed bucket — approximate the smoothness of leaky bucket with less of the implementation complexity, and many production rate limiters (including most managed API gateway products) actually implement a sliding-window variant under the hood rather than a literal bucket. The bucket metaphors are still the clearest way to reason about the two fundamental behaviors — bursty-but-capped versus smoothed-but-delayed — even when the real implementation is a counter and a timestamp.
Choosing between them
Ask what the failure mode actually is. If the goal is protecting your own API from abusive clients while staying friendly to normal bursty usage, token bucket is almost always the right default — it’s what what a REST API client expects when it hits a limit. If the goal is protecting a fragile downstream dependency that genuinely cannot handle spikes, leaky bucket’s queueing and smoothing earns its extra latency.
The takeaway
Token bucket permits bursts up to a cap and limits the long-run average; leaky bucket forces every burst through a fixed-rate drain, trading latency for a perfectly smooth output. Most public-facing APIs want token bucket’s forgiving burst tolerance; infrastructure protecting a fragile downstream wants leaky bucket’s guaranteed ceiling. Pick based on what breaks if a burst gets through unshaped.
Keep reading
The Lycoris Team · · 4 min read What Is the Strangler Fig Pattern?
The strangler fig pattern replaces a legacy system incrementally by routing traffic to new services piece by piece, until nothing old remains.
Chisato · · 4 min read What Are Protocol Buffers? Protobuf Explained
Protocol Buffers (protobuf) are a binary serialization format with strict schemas — smaller and faster to parse than JSON. How it works.
The Lycoris Team · · 5 min read What Is a Distributed Lock?
A distributed lock coordinates exclusive access to a shared resource across multiple processes or machines, preventing race conditions in distributed systems.