Articles

Cloudflare Outages: R2 and Workers Hit in 8-Day Cluster

Cloudflare logged 13 incidents in eight days across R2 storage, Workers KV, and Durable Objects, reviving questions about the web's concentration risk.

Chisato Chisato · · 5 min read
Abstract visualization of a global content-delivery network with glowing connection lines

A string of service disruptions at Cloudflare has put the company’s reliability back under scrutiny. Between August 7 and August 14, 2026, Cloudflare’s own status page recorded 13 separate incidents touching core parts of its developer platform — object storage, key-value, coordination primitives and edge compute — across four continents. None was individually catastrophic. Taken together, they revived a question the industry keeps circling back to: what happens when one company sits in front of a large slice of the web and its foundations wobble repeatedly in the same week?

What broke, and when

The cluster opened on August 7 with a failure in R2, Cloudflare’s S3-compatible object storage, in the Eastern North America (ENAM) region. It closed on August 14 with an availability drop affecting Durable Objects and Workflows, the primitives developers rely on for stateful coordination and long-running orchestration.

In between, the incidents spread across the stack:

  • 503 errors on Magic Transit, the network-layer service that routes and protects customer traffic.
  • Elevated error rates on Workers KV, the globally distributed key-value store many applications use for configuration and session data.
  • Authentication failures on the MCP Server Portal, the entry point for Cloudflare’s Model Context Protocol tooling.
  • Regional 5xx spikes in Kuwait, Bangkok, Jakarta and Dammam, indicating localized network or point-of-presence trouble rather than a single global fault.
  • Intermittent degradation touching Workers AI, the inference layer that runs models at the edge.

The pattern — many discrete, regionally scattered incidents rather than one sweeping outage — is important. It is the difference between a single bad deploy and a run of smaller faults across independent subsystems, and it is harder to wave away as a one-off.

Not the November 2025 event, but not nothing

For scale, none of the August incidents approached the company’s November 18, 2025 global outage, which took down a broad swath of sites at once and became a reference point for how much of the internet depends on a single vendor. This month’s disruptions were narrower and shorter, and most customers experienced them as elevated error rates or regional slowness rather than hard downtime.

But the frequency is the story. Thirteen incidents in eight days is a cadence that shows up in error budgets and on-call rotations even when no single event makes headlines. For teams running production workloads on serverless primitives, a week of intermittent R2 and Workers KV errors can be as disruptive as one clean outage, because retry logic, cache warmth and stateful sessions all degrade in ways that are tedious to diagnose and easy to misattribute.

Racks of network switches and cabling in a data center

The concentration question, again

Cloudflare’s reach is what makes its bad weeks matter. The company operates one of the largest content-delivery and edge networks in the world and, by its own framing, sits in the path of a large fraction of global web traffic. That position is precisely why its developer platform has grown so quickly — a single vendor for CDN, DNS, storage, compute and security is operationally attractive — and precisely why a cluster of incidents ripples outward.

It is a familiar shape. In July, an AWS CloudFront outage rippled across dependent services and underscored how much of the modern web funnels through a handful of edge providers. The lesson each time is the same: consolidation buys simplicity and speed, and it concentrates blast radius. When the vendor you use for storage, coordination and delivery is the same vendor, a bad week for that vendor is a bad week for you, with fewer independent layers to absorb the shock.

The timing is pointed because Cloudflare has been leaning hard into exactly the primitives that stumbled. The company has spent the past year expanding its developer platform and building out agent-oriented tooling, pitching R2, Workers, Durable Objects and Workers AI as a complete backend for AI-era applications. Reliability is the implicit promise underneath that pitch, and a fortnight of R2 and Durable Objects incidents tests it directly.

A rough stretch for developer platforms

Cloudflare was not alone in having a difficult week. Third-party outage trackers logged a cluster of developer-platform disruptions on August 17, including multi-hour incidents at major code-hosting and CI services. The overlap is coincidental rather than causal, but it reinforced a broader unease: the tools developers depend on to ship and run software have felt unusually fragile this month, and each incident lands on teams already stretched by the pace of AI-driven deployment.

For engineering leaders, the practical takeaway is not to abandon a provider after a bad fortnight — every hyperscaler and edge network has incident clusters — but to know precisely where the single points of failure are. Applications that keep all state in one vendor’s key-value store, or all objects in one region of one storage service, inherit that vendor’s worst week with no fallback.

What to watch from here

Two things will tell whether August was noise or signal. The first is Cloudflare’s post-incident transparency: the company has historically published detailed write-ups after major events, and a clear root-cause narrative spanning the R2, Workers KV and Durable Objects incidents would help customers judge whether the faults share a common cause — a network change, a control-plane regression, a capacity issue — or are genuinely independent.

The second is whether the cadence continues. An isolated bad week inside normal variance is very different from the start of a trend. If the incident rate falls back to baseline through late August, the episode reads as an unlucky cluster. If R2 and the coordination primitives keep flaking, it becomes a reliability question that enterprise customers — the ones Cloudflare most wants for its developer platform — will price into their architecture decisions.

What it means

The August incidents are a reminder that reliability is a feature, and that for infrastructure providers it is the feature. Cloudflare’s value proposition — one fast, global platform for delivery, storage and compute — depends on customers trusting that the foundation holds. Thirteen incidents in eight days does not break that trust, but it withdraws from the account.

The winners and losers are subtle. Cloudflare still offers a genuinely compelling, integrated stack, and most teams will (correctly) treat this as a rough patch rather than a reason to re-platform. But the episode strengthens the case for architectural humility: multi-region storage, fallbacks for critical state, and a clear-eyed map of which vendor failures your application cannot survive. The teams that fared best this month were the ones who had already asked that question — and had answers that did not begin and end with a single provider.

If you build on Cloudflare’s edge — including by deploying static sites to Cloudflare Pages or running APIs on Workers — the practical move is not fear but diligence: read the post-incident reports when they land, know your dependencies, and design for the assumption that any single layer can have a bad week. Because, eventually, every one of them does.

Chisato Chisato · · 5 min read

What Is a VPC Endpoint?

A VPC endpoint gives a private network direct access to a cloud service without routing traffic through the public internet or a NAT gateway.

#Cloud #Networking #Infrastructure
Chisato Chisato · · 4 min read

Azure Outage Takes Down ChatGPT, Claude, and Grok

An Azure East US ingress failure knocked ChatGPT, Claude, and Grok offline at once on Sept 3 — the first outage to visibly take down three rival AI labs together.

#Azure #Cloud #Infrastructure
Chisato Chisato · · 5 min read

Alibaba Cloud Launches First Brazil Region

Alibaba Cloud opened its first South American region in São Paulo, with two data centers and planned agentic AI services, part of a $53B infrastructure push.

#Cloud #Alibaba #AI