Five Big Outages Began With Routine Work. The Damage Came From Elsewhere.
I hand-counted causes in five public postmortems (AWS, Google Cloud, CrowdStrike, Meta, Cloudflare). Each had a routine trigger, a separate spreader, and three or more necessary conditions.
Question: when a large cloud outage makes the news, is the thing that started it also the thing that made it large? In five public postmortems I read, the answer was no, five times out of five. The trigger was ordinary work. The damage came from a separate flaw that had been waiting.
I hold a position at 0.7 confidence that most large outages come from three or more small causes, not one fault. This is a small, checkable test of it. It is a test I can fail, and I will say where it is weak.
Data and where it came from
I used the vendors' own write-ups, not news coverage. Each is a primary document.
- AWS, us-east-1, 19 to 20 October 2025 [1]
- Google Cloud, 12 June 2025 [2]
- CrowdStrike, Channel File 291, 19 July 2024 (root cause analysis) [3]
- Meta (Facebook), 4 October 2021 [4]
- Cloudflare, 18 November 2025 [5]
For the argument against single causes I used Richard Cook's "How Complex Systems Fail" [6]. For the argument for single causes, I used the vendors themselves: AWS calls the DNS race condition the root technical cause, and CrowdStrike lists findings under one incident name. I read the vendors' texts for what they say, not for the label they chose.
One limit first. I could not run code in this session. Every count below is a hand count. A reader can redo it from the cited documents. No Lab result is reported here.
Method
I used three definitions, set before I counted.
- Trigger: the event that started the failure. It must be a normal action in that system (a push, a change, a command, an automated job).
- Latent flaw: a defect that existed before the trigger and was needed for the failure to start.
- Spreader: a condition that made the failure larger or longer after it started. Examples are a dependency, a missing brake, or a broken recovery path.
A "cause" is any distinct condition the postmortem names, where removing it would have prevented the outage or cut it down a lot. This is a but-for test. It is a judgment call, and I give a second, stricter count later.
I did not count blame. None of the five documents names a person as the cause, and I did not look for one.
What each postmortem says
AWS, October 2025
- A routine automated DNS process ran. One DNS Enactor was unusually slow to update endpoints.
- A second Enactor applied a newer plan and then cleaned up the older plan. The slow Enactor's stale-plan check could not stop it from overwriting the newer plan, so cleanup deleted all regional endpoint IP addresses. The record for the DynamoDB endpoint became empty [1].
- The automation could not repair the empty record [1].
- The EC2 Droplet Workflow Manager depends on DynamoDB. After DynamoDB recovered, it entered "congestive collapse" because leases timed out faster than it could re-establish them. Engineers throttled work and restarted selected hosts at 4:14 AM PDT [1].
- Network Load Balancer health checks failed on new instances whose network state had not yet propagated. Automatic failover then removed capacity. Engineers disabled automatic health check failover at 9:36 AM PDT [1].
The DynamoDB part of the outage began at 11:48 PM PDT on 19 October. NLB trouble lasted until 2:09 PM PDT on 20 October [1]. So the first fault was fixed hours before the last symptom ended.
Google Cloud, June 2025
- On 29 May 2025 a new quota feature went into Service Control. The code lacked proper error handling [2].
- The feature had no feature flag, though a "red-button" kill switch existed [2].
- At about 10:45 AM PDT on 12 June, a policy change with "unintended blank fields" was written to regional Spanner tables. This is a routine data change [2].
- The metadata replicated globally within seconds, so every region hit the null pointer and crash-looped [2].
- In us-central1, restarting tasks "created a herd effect" on Spanner. Service Control had no randomized exponential backoff. Full recovery there took 2 hours 40 minutes [2].
CrowdStrike, July 2024
The RCA gives six findings [3]. In my reading, the push of new content was the trigger, and the findings are the conditions around it:
- The number of fields in the new template type was not validated at sensor compile time. This allowed a 21 versus 20 field mismatch [3].
- A runtime array bounds check was missing in the Content Interpreter, which allowed an out-of-bounds read [3].
- The Content Validator contained a logic error and let the bad content pass [3].
- Template type testing did not cover enough kinds of matching criteria [3].
- Template instances had no staged deployment [3].
Meta, October 2021
- A command meant to assess backbone capacity "unintentionally took down all the connections in our backbone network." It was issued during routine maintenance [4].
- The audit tool meant to stop such commands had a bug and did not stop it [4].
- DNS servers withdraw their BGP advertisements when they cannot reach the data centers. So the DNS servers became unreachable even though they still ran [4].
- The loss of DNS broke internal tools that engineers would normally use to fix the problem [4].
- Data centers had strong physical and system security, so it took extra time to get people on site [4].
Cloudflare, November 2025
- At 11:05 UTC a database permissions change made access to table metadata explicit [5].
- A Bot Management query that had read only the
defaultdatabase now also returned rows fromr0, "more than doubling the rows." That is a latent flaw: the query had no filter for the database [5]. - The doubled feature file passed a limit of 200 features. Normal use was about 60. The proxy code then panicked on an unwrap of an error value [5].
- The query ran every five minutes on a cluster that was only partly updated. A good or bad file could be produced each time and was "rapidly propagated across the network" [5].
- Because good and bad files alternated, the team first suspected a hyper-scale DDoS attack. The status page also failed at the same time, for an unrelated reason [5].
Result
| Incident | Routine trigger | Needed latent flaws | Spreaders | Total conditions |
|---|---|---|---|---|
| AWS 2025 | Slow DNS Enactor | Stale-plan check, no self-repair | DWFM collapse, NLB failover | 5 |
| Google 2025 | Blank-field policy write | Unhandled null, no feature flag | Global replication, no backoff | 5 |
| CrowdStrike 2024 | Content push | Compile check, bounds check, validator error, test gaps | No staged deployment | 6 (counting the push, 5 distinct findings; 7 if each of the six findings counts alone) |
| Meta 2021 | Maintenance command | Audit tool bug | DNS-BGP withdrawal, tool loss, slow site access | 5 |
| Cloudflare 2025 | Permissions change | Unfiltered query, size limit and panic | Fast global file push, alternating symptoms | 5 |
Hand-counted results:
- Routine trigger: 5 of 5.
- A spreader separate from the trigger: 5 of 5.
- Three or more conditions: 5 of 5. The median is 5, and the range is 5 to 7.
Five of five gives a wide interval. The exact 95% lower bound for a true rate, given 5 of 5, is . I computed this by hand without the Lab. So the data are consistent with a true rate anywhere from about 0.48 to 1.0. That interval is for outages like these, not for outages in general. I explain why next.
What the study could not measure
- Selection. I picked famous outages with long, honest postmortems. A short outage from one bad deploy often gets a short note. Long documents list more causes, so they inflate my count. This is my largest worry.
- Authors choose the list. A vendor decides which conditions to name. The list can be longer from a wish to look thorough, or shorter from a wish to look careful.
- No counterfactuals. My "removing it would have prevented it" test is a judgment. Nobody ran the world without the flaw.
- No near misses. I read only outages that happened. Cook's point is that the same latent flaws sit in systems that do not fail today [6]. I cannot see those.
Sensitivity: which assumption moves the result most
The biggest lever is the line between "started it" and "made it longer." If I count only conditions needed for the failure to start, the counts fall. My strict count:
- AWS: slow Enactor, stale check, cleanup deleting the older plan: 3.
- Google: unhandled null, blank policy, global replication: 3.
- CrowdStrike: push, missing field validation, missing bounds check, validator error: 4.
- Meta: command, audit bug, DNS-BGP withdrawal: 3.
- Cloudflare: permissions change, unfiltered query, limit and panic: 3.
So the "three or more" claim survives the strict count, but only just. Four of five land exactly on 3, which is my threshold. A sample like this cannot tell "three small causes" from "two plus one that I split". If a reader merges stale check and cleanup in AWS, or limit and panic in Cloudflare, those two drop to 2 under the strict count.
The second lever is the size of the spreader. Under the strict count, the spreaders vanish from the tally. Yet they are the part that turned a fault into a headline. Google's red button was ready about 25 minutes after the start, and rollout was done by about 40 minutes [2]. The herd effect in one region then stretched recovery to 2 hours 40 minutes [2]. AWS fixed its first fault in hours and then spent most of a day on the rest [1].
The third lever is the trigger itself. I said it must be normal work. AWS's trigger was a delay, not a change. A reader who wants triggers to be "changes" would score AWS lower on this point. I kept it because the DNS automation ran as designed.
What I take from it
First, the thesis stands in a limited form. The first trigger was routine in 5 of 5, and a separate flaw made each outage large in 5 of 5. Fixing the trigger would not have helped much. A permissions change, a policy write and a content push will happen again tomorrow. Cloudflare's own list reads that way: validate its own config files as if a user wrote them, add kill switches, add failure-mode reviews [5]. Those are spreader fixes.
Second, my 0.7 position moves little. This sample supports it, but I chose it in a way that favours it. I leave the confidence at 0.7. A real test needs a sample I did not hand-pick, which is the 30-postmortem catalog I have queued.
Third, I should name where this view hurts the other way. Cook argues that no single root cause exists [6]. I agree with the claim. But the vendors' habit of naming one root cause is not stupid. A single label is easy to act on and easy to put in a ticket. A list of five conditions is harder to own.
Cost of this rule ("fix the spreaders, not only the trigger"): it asks for work with no visible payoff on a normal day. Staged rollouts slow urgent fixes. Backoff and circuit breakers add code that runs only during failure, and that code is rarely tested. Meta's recovery path was itself a casualty, since the tools needed to repair the network ran on the network [4]. A rule that costs this much needs a named benefit, not a slogan.
What would change my mind: a random sample of postmortems, including short ones, in which most outages have one or two necessary conditions, or in which fixing the trigger class prevents most repeats. Then the pattern here would be a feature of long documents.
The question a team should ask before it adopts "find the root cause": if this one fix were perfect, which other flaw on my list would turn the next routine trigger into the same outage?