TCP Congestion Control Explained: Slow Start to BBR
TCP congestion control decides how fast a sender transmits without overloading the network. Slow start, AIMD, CUBIC, and BBR explained.
TCP congestion control is the set of algorithms a TCP sender uses to decide how much data it can have in flight at once, so that it uses available bandwidth without overwhelming the routers and links between it and the receiver. The sender has no direct view of the network. It infers congestion from signals such as lost packets and rising round-trip times, and adjusts a variable called the congestion window in response. Every TCP connection you open, including every HTTPS page load, runs one of these algorithms.
Why congestion control exists
Routers have finite buffers. If senders collectively push more traffic than a link can carry, queues fill, packets are dropped, senders retransmit, and the retransmissions add even more load. Without a feedback mechanism, a network can collapse into a state where links are busy but almost no useful data gets through. Congestion control was added to TCP to prevent exactly that, and it remains the reason the internet can be shared fairly among billions of connections with no central coordinator.
It is distinct from flow control. Flow control protects the receiver: the receiver advertises how much buffer space it has, and the sender never exceeds it. Congestion control protects the network. At any moment, a sender may transmit no more than the smaller of the receiver’s window and its own congestion window.
The congestion window
The congestion window (cwnd) is the amount of unacknowledged data the sender allows itself to have in flight. Throughput is roughly cwnd / RTT, so a connection with a 100 ms round trip and a 1 MB window moves about 10 MB per second. The algorithms below are all just different rules for growing and shrinking cwnd.
Slow start
A new connection does not know how much capacity is available, so it starts with a small window of a few segments. For every acknowledgment received, it increases cwnd by one segment. Because each round trip’s worth of ACKs grows the window, the window roughly doubles every round trip. “Slow” start is therefore exponential; it is only slow relative to blasting at full speed immediately.
Slow start continues until either a loss is detected or cwnd reaches the slow-start threshold (ssthresh), after which the connection switches to a gentler growth mode.
Slow start is why short-lived connections rarely reach full bandwidth, and why reusing connections, as HTTP/2 and HTTP/3 do, matters so much for web performance. A fresh connection spends its first several round trips ramping up. It also explains why time to first byte and the size of the first few kilobytes of a page are so sensitive to latency.
Congestion avoidance and AIMD
Above ssthresh, classic TCP uses additive increase, multiplicative decrease (AIMD):
- Additive increase: grow
cwndby about one segment per round trip, probing gently for more bandwidth. - Multiplicative decrease: when loss is detected, cut
cwndsharply, traditionally by half.
Plotted over time, this produces TCP’s familiar sawtooth. AIMD has a useful property: when several flows share a bottleneck, it drives them toward an equal share of capacity.
Detecting loss
TCP learns about loss in two ways:
- Duplicate ACKs. If the receiver gets packets out of order, it keeps acknowledging the last in-order byte. Three duplicate ACKs strongly suggest one packet was lost while later ones arrived. The sender retransmits immediately (fast retransmit) and halves the window rather than starting over (fast recovery).
- Retransmission timeout. If no ACKs arrive at all, the sender waits for a timer to expire, assumes severe congestion, and drops
cwndback to its initial small value, restarting slow start. Timeouts are expensive, which is one reason packet loss shows up so clearly in tail latency.
The major algorithms
Reno and NewReno
The classic loss-based algorithms implementing slow start, AIMD, fast retransmit, and fast recovery. NewReno improved recovery when several packets are lost from the same window. They work well on modest links but grow too slowly on paths with high bandwidth and long round trips, where adding one segment per RTT can take a very long time to fill the pipe.
CUBIC
CUBIC replaces linear growth with a cubic function of time since the last loss. After a reduction, the window grows quickly back toward its previous maximum, flattens out near it, then accelerates again to probe beyond. Because growth depends on elapsed time rather than ACK arrival, it is fairer between flows with different round-trip times. CUBIC has long been the default in Linux and other major operating systems.
BBR
BBR (Bottleneck Bandwidth and Round-trip propagation time) takes a different approach. Instead of treating loss as the signal, it builds a model of the path: the bottleneck bandwidth, estimated from delivery rate, and the minimum round-trip time. It then paces packets to send at roughly the bottleneck rate while keeping only about a bandwidth-delay product of data in flight. The aim is to fill the pipe without filling router buffers.
Loss-based vs model-based at a glance
| Loss-based (Reno, CUBIC) | Model-based (BBR) | |
|---|---|---|
| Primary signal | Packet loss | Measured bandwidth and min RTT |
| Behaviour with deep buffers | Fills them, raising latency | Tries to keep queues short |
| Behaviour with random loss | Backs off unnecessarily | Largely unaffected |
| Sending style | Bursty, window-driven | Paced |
| Fairness concerns | Well understood | Can compete unevenly with loss-based flows |
Bufferbloat and why it matters
Loss-based algorithms keep growing the window until a buffer overflows. If routers have very large buffers, those buffers fill before any loss occurs, and every packet waits in a long queue. Throughput looks fine, but latency balloons, which ruins interactive traffic such as video calls, games, and WebSocket apps sharing the link. This is called bufferbloat. Two broad fixes exist: delay-aware algorithms like BBR on the sender, and active queue management on routers, which drops or marks packets early to signal congestion before queues grow long. ECN (Explicit Congestion Notification) lets routers mark packets instead of dropping them, giving senders the signal without the loss.
Congestion control beyond TCP
QUIC, the transport underneath HTTP/3, implements congestion control in user space rather than in the kernel, typically using the same families of algorithms. That makes it easier to deploy improvements without operating-system upgrades. Our HTTP/2 vs HTTP/3 comparison covers QUIC’s other changes, and TCP vs UDP explains why QUIC builds on UDP.
The takeaway
TCP congestion control lets every sender share the network without a central controller by adjusting a congestion window based on feedback. Slow start ramps up quickly, AIMD probes gently and backs off hard, CUBIC scales that idea to fast long-distance links, and BBR models the path to avoid filling buffers at all. For application developers, the practical lessons are to reuse connections, keep early payloads small, and remember that loss and queueing, not just bandwidth, shape how fast things feel.
Keep reading
The Lycoris Team · · 6 min read Copy-on-Write Explained: How Lazy Copying Works
Copy-on-write shares data until someone modifies it, then copies only what changed. How CoW powers fork(), snapshots, containers, and immutable data.
The Lycoris Team · · 5 min read Head-of-Line Blocking Explained: HTTP/1.1, HTTP/2, QUIC
Head-of-line blocking is when one stalled item holds up everything queued behind it. How it affects HTTP/1.1, HTTP/2 over TCP, and how QUIC fixes it.
Chisato · · 4 min read Thermal Interface Materials Explained
Thermal interface material fills microscopic gaps between a chip and its heatsink so heat can actually transfer to the cooler.