Articles

What Is the Bulkhead Pattern? Isolating Failures

The bulkhead pattern isolates resources per dependency so one failing service can't exhaust the pools that healthy services depend on.

Chisato Chisato · · 4 min read
Server racks with bundled network cables

The bulkhead pattern partitions a system’s resources — thread pools, connection pools, or concurrent request slots — into isolated groups per dependency, so that one failing or slow dependency can’t exhaust the resources healthy dependencies need. It trades some efficiency (idle capacity sits unused in one partition while another is saturated) for a hard limit on how far a single failure can spread.

Where the name comes from

The pattern borrows its name from ship design: a hull is divided into watertight compartments, so a breach in one section floods only that compartment instead of sinking the whole ship. Software failures behave the same way when resources are shared. If a single thread pool serves calls to three downstream services and one of them starts responding slowly, every thread can end up blocked waiting on that one slow dependency — starving the other two services of any capacity at all, even though they’re healthy.

What it isolates in practice

The most common target is the pool of resources used to call out to a dependency:

  • Thread pools — a dedicated pool per downstream service, so a hang in service A can’t consume the threads service B’s calls need.
  • Connection pools — separate database or HTTP connection pools per dependency, preventing one saturated connection pool from starving another.
  • Concurrency limits (semaphores) — capping how many concurrent calls to a given dependency are allowed, independent of thread allocation, which is cheaper than dedicating whole thread pools when the calls are short-lived.

The partitioning can also happen at a coarser level — separate deployments, separate nodes, or separate Kubernetes namespaces per tenant or workload class — so that one noisy tenant can’t degrade service for everyone else sharing the cluster.

Bulkhead vs circuit breaker

The two patterns are often deployed together but solve different problems.

BulkheadCircuit breaker
Problem it solvesResource exhaustion spreading across dependenciesRepeatedly calling a dependency that’s already failing
MechanismIsolated pools/limits per dependencyTrips open after a failure threshold, stops calling
Failure containmentLimits blast radius of one slow/failing callStops wasted calls once failure is detected
RecoveryAutomatic — capacity frees as calls completeRequires a half-open probe state to test recovery

A circuit breaker decides whether to keep calling a dependency at all; a bulkhead decides how much of the system’s total capacity that dependency is allowed to consume while it’s being called. Using only a circuit breaker still leaves the door open for one dependency’s calls to occupy every available thread before enough failures accumulate to trip the breaker — which is exactly the scenario a bulkhead prevents.

Sizing the partitions

Bulkheads only help if the partitions are sized sensibly. Too generous, and a struggling dependency can still consume enough shared infrastructure (CPU, memory, outbound network bandwidth) to degrade everything else even with its own thread pool. Too conservative, and legitimate traffic gets rejected during normal spikes because its partition is undersized relative to actual demand.

A reasonable starting point is to size each partition around the dependency’s typical latency and expected concurrent call volume, then load-test the failure scenario deliberately — call the isolated dependency with an artificially slow or failing response and confirm that calls to other partitions are unaffected. This is exactly the kind of failure injection that chaos engineering practices formalize: proving isolation holds under real degradation rather than assuming it from the architecture diagram.

Bulkheads and backpressure

When a partition is exhausted, the system needs a decision for what happens next: queue the request, reject it immediately, or degrade to a fallback response. Queuing without a bound just delays the exhaustion problem and adds latency; an unbounded queue behind a bulkhead can still lead to the same resource exhaustion it was meant to prevent, just further downstream. Pairing bulkheads with explicit backpressure — rejecting or shedding load once a partition’s queue passes a bound, rather than letting it grow indefinitely — keeps the isolation meaningful under sustained load rather than just delaying the failure.

When to reach for it

Bulkheads earn their complexity in systems that call multiple independent external dependencies from a shared process — a checkout service that calls payments, inventory, and a recommendations service, for example, where a recommendations outage has no business taking down checkout. They matter less in a strict microservices architecture where each service already runs in its own process and node, since the container boundary itself provides much of the isolation a bulkhead would add inside a monolith. Multi-tenant systems are the other common case: isolating each tenant’s resource consumption, whether via per-tenant connection pools or Kubernetes HPA/VPA-managed resource quotas, keeps one tenant’s traffic spike from degrading service for everyone sharing the platform.

The takeaway

The bulkhead pattern isolates the resources each dependency uses — thread pools, connection pools, or concurrency limits — so a failure in one dependency can’t consume the capacity another dependency needs to keep working. It’s most valuable in processes that call several independent, unreliable dependencies from shared infrastructure, and it pairs naturally with circuit breakers (which stop calling a failing dependency) and explicit backpressure (which bounds what happens once a partition fills up). Size partitions from real latency and load data, then verify the isolation holds by deliberately injecting a failure rather than trusting the diagram.

The Lycoris Team The Lycoris Team · · 4 min read

What Is the Strangler Fig Pattern?

The strangler fig pattern replaces a legacy system incrementally by routing traffic to new services piece by piece, until nothing old remains.

#Backend #Distributed Systems #Software Architecture
Chisato Chisato · · 4 min read

What Are Vector Clocks? Ordering Distributed Events

A vector clock is a per-node counter array that lets distributed systems tell whether one event happened before another, without a shared clock.

#Distributed Systems #Databases #Cloud