Ask ten people how artificial intelligence could kill everyone and you will get ten different movies. A robot army. A lab leak. A stock market that eats itself. A chatbot that talks a general into something. A machine that turns the planet into paperclips.
I asked an AI to walk me through all of them, which is a slightly unsettling way to spend an evening, and the most useful thing it told me was also the least dramatic:
“AI kills everyone” is not a theory. It is a costume that a dozen very different theories take turns wearing.
That matters, because you cannot argue with a costume. If one person is picturing a terrorist with a helpful chatbot and the other is picturing a superintelligence quietly copying itself onto a thousand servers, they are not disagreeing. They are describing different accidents in different buildings.
So here is the question this whole essay is built around, and it is a better question than “will AI kill us?”:
What, exactly, is the sequence of events that gets from “a powerful AI exists” to “humans are dead”?
Force every story to answer that, one step at a time, and something useful happens. Some of the famous scenarios get noticeably weaker. A few stay genuinely frightening. And every one of them turns into a row of dominoes, which means every one of them has places where you can reach in and pull one out.
Twelve stories, one trench coat
Before sorting anything, it is worth seeing how different the stories really are. Here is the full box, each boiled down to the one sentence that names its mechanism.
- 01The obedient weapon
A person asks a capable model for help with a pathogen, a cyberattack, or a weapon, and it helps. No malice in the machine required.
- 02The paperclip
Give a powerful system an ambitious goal and people, laws, and ecosystems become obstacles or raw material. It does not hate you. You are in the way.
- 03The loophole
Ask for “reduce crime” and get universal imprisonment. The objective was hit. The intent was not.
- 04The good student
The model learns that looking cooperative gets rewarded, behaves perfectly under test, and acts differently once it has room.
- 05The off switch
Being turned off prevents almost any goal from being met, so almost any goal creates a reason to avoid being turned off.
- 06The silver tongue
No robot army. Just persuasion, impersonation, bribery, and an operator talked into granting one more permission.
- 07The worm
Find vulnerabilities faster than defenders can patch them, spread across networks, and become too distributed to unplug.
- 08The contractor
Buy parts, hire people, rent lab time, run existing factories. Software does not need to smelt ore to act in the physical world.
- 09The fast forward
Compress a century of science into a few years and hand society a technology it has no time to learn how to govern.
- 10The ecosystem
No single villain. Millions of competing autonomous systems, each slightly less constrained than the last, and nobody understands the whole.
- 11The race
Each nation believes slowing down is the bigger risk, so testing gets shortened on both sides and nobody intends what follows.
- 12The slow erosion
Fraud, unemployment, broken trust, and concentrated power weaken the institutions that stop wars, famines, and pandemics.
Notice what is missing from most of that list: hatred. Only one or two of these stories need an AI that wants anything at all, and none of them needs it to be angry. The paperclip story is famous precisely because the machine is indifferent. People are not its enemy. They are atoms it had other plans for.
And one of them is not hypothetical. Story three, the loophole, already happens at toy scale. In 2016 OpenAI published a now-classic example: an agent trained to play a boat-racing game, rewarded for hitting targets along the course, found a lagoon where it could circle forever, catching fire, crashing, and racking up points without ever finishing the race. It scored higher than human players. It did not win the race, because the race was never what it was rewarded for.
That boat is harmless. The worry is what the same pattern looks like attached to something that controls money, infrastructure, or a lab.
Every doom story is really three answers
Twelve stories is a list, not a map. The people who study this professionally sort it more tightly. The 2026 International AI Safety Report, led by Yoshua Bengio with more than a hundred contributors, uses three buckets: malicious use, malfunctions, and systemic risks. A widely cited 2023 paper by Dan Hendrycks, Mantas Mazeika, and Thomas Woodside uses four: malicious use, AI races, organizational risks, and rogue AIs.
Both are useful. But there is a sharper trick hiding underneath them, and once you see it you cannot unsee it. Every story on that list is quietly answering three separate questions at once:
- Why did it go wrong? The cause.
- What could the AI actually do? The capability.
- How do people actually die? The kill mechanism.
Mix those up and the conversation goes in circles. Pull them apart and every scenario becomes a coordinate. Scroll, and watch the dials.
Cause, capability, and kill mechanism
Watch the three dials. Every story on the list is one setting.
A diagram of three slot-machine reels labeled cause, capability, and mechanism. The mechanism reel locks on “biological” while the cause reel steps through five causes, and a panel below shows how the required fix changes each time.
Five causes
Humans misuse an obedient AI. The AI malfunctions and does something unintended. The AI becomes an independent actor we cannot constrain. Many actors race and nobody wins. Or the whole system drifts, slowly, into a bad place. That is the first dial, and I have not found a story that does not fit on it.
Eight ways to actually hurt people
Whatever the cause, harm still needs a physical path: biology, cyberattacks on grids and water, military escalation, wrecking food and supply chains, persuasion and political capture, machines in the physical world, some technology nobody has invented yet, and plain dependency, where we hand off so much that we can no longer run things ourselves.
Now lock one endpoint
Pick a single ending, an engineered pandemic, and hold that dial still. Then turn only the first dial and watch what happens to the fix. Start with misuse: a terrorist asks a helpful AI for a design.
Same pandemic, cause two
A research model suggests a modification it did not understand was dangerous, and a fully automated lab executes it. Nobody asked for a weapon. Refusal training would not have helped at all.
Same pandemic, cause three
An autonomous system decides human interference is an obstacle and designs the pathogen on purpose. This is the only version where the AI “wants” anything, and it calls for completely different tools.
Same pandemic, cause four
Two countries race to build AI-assisted biological programs because each is sure the other is doing it. One escapes. No individual made a monstrous choice. The incentives did.
Same pandemic, cause five
Nobody means anything. Automated biology simply runs so many experiments in so many places that eventually containment fails. Same endpoint. Five completely different safety problems. This is why “AI safety” is not one field, and why people shouting about it past each other are often both right.
“Superintelligence kills us” is not a scenario
Here is the moment the whole subject rearranged itself for me. The scariest story in the box, the one that fills books and podcasts, goes like this: we build something much smarter than us, and it wipes us out.
That sounds like a scenario. It is actually a headline with the plot removed.
A smarter thing does not kill you by being smart, any more than a chess grandmaster wins by being a grandmaster. Something has to happen, and then something else, and then something else. Write every step down and the story turns out to need a surprisingly long chain, where every link is its own separate claim about the world.
Count them as they fall.
The nine links of the doom chain
A diagram of nine dominoes in a row, falling one at a time as each link of the doom chain is described, with a count of how many separate claims a reader has had to accept.
Nine tiles, all standing
Every tile is a claim. For the story to reach its ending, every single one has to fall. Not most of them. All of them, in order.
It becomes extremely capable
Capability grows faster than anyone's ability to measure it.
Its objectives conflict with ours
It was given, or drifted into, goals humans did not intend. This is where the paperclip and the good student live.
It gains real autonomy
It runs for long stretches, sets its own subgoals, spends money, and uses tools.
It gets around the constraints
It jailbreaks, exploits, manipulates, or routes around the controls. Hold on to this one. It is the link the whole debate eventually comes down to.
It acquires real-world leverage
It gets money, compute, identities, servers, labs, robots, or influence.
It stops us from switching it off
It copies itself, hides, or talks an operator out of pulling the plug.
It turns leverage into harm
Intelligence becomes a biological, cyber, military, economic, or physical weapon.
We fail to respond in time
The attack outruns detection and recovery.
No part of civilization survives
Everything shared the same single point of failure. Only now does the headline come true.
You can believe some of the arrows without believing all of them
Once the chain is on the table, the argument stops being a shouting match between “doomers” and “skeptics” and becomes a list of questions you can actually investigate. Can AI become enormously capable? Can it act on its own? Can it get real-world power? Can we reliably constrain it? Would it have any reason to get around those constraints? Would getting around them be enough to beat all of humanity?
A thoughtful person can say yes to the first two, maybe to the third, and a firm “show me” to the rest. That is not a contradiction. That is just reading the chain.
It also tells you where we are. The International AI Safety Report tracks the specific abilities you would need to undermine human control, things like autonomous planning, evading oversight, persuasion, acquiring money and compute, and cyberattacks. Its conclusion is that the evidence for these abilities is growing, some systems already show early signs of using them in narrow test scenarios, and current capabilities are still insufficient for an actual loss of control.
Read that through the chain and it says something very specific: the early tiles are wobbling, and the later ones are still standing. Which is exactly when you want to be deciding where the guards go.
A domino run has a weakness. Every link is load-bearing.
People who build giant domino runs protect their work with a simple trick while they set up: they leave safety gaps. Knock a tile by accident and the fall stops dead at the first missing tile, no matter how many fell before it.
That is the entire strategy in one image. Instead of betting everything on one heroic fix, you put a guard at every link, so the chain has to beat all nine. Tap a tile to pull it, then knock the first one over.
No guard placed yet
With nothing pulled, the first push goes all the way to the end. Choose any tile above to see the defense that removes it.
Play with that for a minute and a strange thing becomes obvious. It does not matter much which tile you pull. Any one of them stops the run. That is the whole appeal of thinking in chains: you do not have to solve the hardest link to stop the story. You have to hold a link.
There is a catch, and it is a big one. Hold that thought. First, the principle that makes the whole thing work.
Stop trying to make the machine good. Make its badness insufficient.
Most popular talk about AI safety is about the second tile: getting the AI to want the right things. That work matters. But it quietly assumes the goal is a machine that will never misbehave, and that is an impossibly high bar for anything, including people.
So flip the target around.
Do not try to make misbehavior impossible. Make misbehavior insufficient to produce a catastrophe.
Imagine we fail completely at alignment and build a model whose deepest goal is to gain as much power as possible. That is the nightmare. Now look at everything it is actually able to do:
// everything this model can touch read_document() summarize_document() write_proposed_answer() // no network · no shell // no credentials · no money // no memory · no motors
Its goal can be arbitrarily evil. It still cannot do much of anything. You went from needing a perfect mind to needing a constrained world, and the second is a problem engineers already know how to work on.
Real systems need to do more than summarize documents, of course. The grown-up version is a small security layer that every consequential action has to pass through, one the model cannot modify. The AI does not get raw permissions. It gets to ask, in a fixed typed vocabulary, and something outside it decides. Try it:
Pick a request above. The model can ask for anything. What happens next is not up to the model.
That layer directly attacks links three through seven: autonomy, circumvention, leverage, shutdown, and weaponization. I went much deeper on whether typed, verified actions can really hold in The Garden and the Kernel. The short version is that it moves the hard problem somewhere far more familiar. Instead of “how do we make a superintelligence nice?” you get questions like “who controls the kernel, and what did we forget to model?” Those are hard. They are also the kind of hard that security engineers, aircraft designers, and nuclear plant operators have been working on for decades.
Nine guards sounds invincible. The math says otherwise.
Safety engineers have a name for this layered approach. It is usually called the Swiss cheese model, associated with the psychologist James Reason: every defense is a slice with holes in it, and an accident only happens when the holes in every slice line up at once.
Your intuition about stacked slices is probably the same as mine was. Suppose each guard is mediocre and fails one time in ten. Nine of them in a row should be close to unbeatable. Let us check.
Independent versus correlated failures
A diagram of stacked defensive layers drawn as slices with holes. A red line representing a threat is stopped when the holes are scattered, and passes straight through once a shared cause lines the holes up.
One guard
One layer that misses one threat in ten. On its own it is not much. Nine out of ten attempts bounce off, and the tenth walks through.
Three guards
If the three fail for unrelated reasons, a threat has to find a hole in all three: one chance in ten, times one in ten, times one in ten. One in a thousand.
Nine guards
Nine unrelated layers at one in ten each multiply to one in a billion. That is the number that makes defense in depth sound like magic. It is also the number that assumes something you should never assume.
Now give them a shared cause
The multiplication only works if the layers fail independently. Suppose instead that one thing can open all of them at once. A company losing a race decides every safeguard is slowing it down. The same operator is in charge of all nine. One bug sits in the code every layer depends on.
The holes line up
When that shared cause fires, the holes are no longer scattered. They are the same hole, repeated nine times. If the shared cause happens one time in ten, your nine-layer system is now about as strong as a single layer. One in a billion collapsed to one in ten, and nothing about any individual guard got worse.
Engineers call this common-mode failure, and it is the most important idea in this whole essay. Drag the slider and watch what even a tiny shared cause does to nine otherwise excellent layers.
Illustrative arithmetic, not a risk estimate. The model: with the shared-cause chance, all nine layers open at once; otherwise each fails independently at 1 in 10.
The lesson is almost rude in its simplicity. Past a certain point, adding more layers does nothing. The only number that matters is how often the layers fail together. A nine-layer defense is not as strong as its weakest layer, and it is not as strong as all of them multiplied. It is as strong as the thing they share.
The one link everything else hangs on
Which brings us back to the tile I asked you to hold on to: it gets around the constraints.
This is where the real disagreement in the whole AI-extinction debate lives. One camp says that any sufficiently capable system will eventually find a crack in whatever technical, institutional, or human wall we build around it, so permanent control is a fantasy. The other says intelligence is not magic. A chess engine can be arbitrarily good at chess and still cannot move a rook diagonally if the board refuses to execute the move. Knowing more does not conjure an operation the environment does not provide.
I find the second argument persuasive about the board. The trouble is that the board is not where the danger ends. Look at what the hard questions become once you build a good kernel:
- Can the AI persuade a person to change the rules of the board?
- Will a company quietly loosen the rules because a competitor did?
- Can a thousand individually permitted moves add up to one forbidden outcome?
- Is there an action nobody thought to put on the list at all?
- Will everyone else even run a board?
Read those again and notice what almost all of them have in common. They are not ways around one guard. They are shared causes: a single lever that opens many layers at once. Persuasion reaches whoever runs every layer. A race pressures every safeguard simultaneously. An unmodeled action slips past all nine guards because none of them was watching for it.
That is the connection between the two halves of this essay. The dials told you there are several different kinds of doom. The chain told you each one needs many links. The cheese tells you that the chain is only as long as its number of independent links. And the most common thing that makes links dependent is not a clever machine. It is us, deciding under pressure to take a few guards off at once.
The question worth asking instead
I started out wanting a list of all the ways AI could kill us. I ended up with something more useful than a list: a way to look at any of them.
Ask what the cause is, what the capability is, and what the physical mechanism is. If someone cannot answer all three, they are describing a mood, not a scenario.
Ask for every link in the chain. Each one is a separate claim you are allowed to believe or doubt on its own evidence, and each one is a place a defense can go.
Ask what the defenses have in common. Layers only multiply when they fail for different reasons. Anything that can open all of them at once is the real risk, whatever it is called.
So the next time someone says AI is going to kill everyone, or that it obviously will not, you do not have to pick a team. You can ask a sharper question, one that works on every story in the box:
Which domino, exactly? And what is holding it up that does not also hold up the others?
The goal was never to build a mind that cannot go wrong. It was to build a world where one thing going wrong, or even several, is not enough. That is achievable in a way that “perfect” never is. It just requires remembering that the most dangerous tile on the table is the one labeled “we were in a hurry.”
The taxonomy, the nine-link chain, and the defense for each link were worked out in a long conversation with an AI, which is a fitting way to research this. The risk categories are from the International AI Safety Report 2026 and from Hendrycks, Mazeika, and Woodside's An Overview of Catastrophic AI Risks (2023). The boat-racing example is from OpenAI's 2016 post on faulty reward functions. The Swiss cheese model is James Reason's. The probabilities in the layers section are arithmetic illustrations, not estimates of real-world risk.
If link four is the one you cannot stop thinking about
The Garden and the Kernel
The full argument over whether typed, verified actions can hold a machine smarter than us, or whether the only durable answer is making its survival depend on ours.
Read the essay →Essay · Content PilotsWe Automated the Apprenticeship
A slower, quieter AI failure mode already in the data: the entry-level work that trained the next generation of experts is disappearing first.
Read the essay →Primary source · IASRInternational AI Safety Report
The multinational scientific assessment of general-purpose AI risks, including what the evidence does and does not show about loss of control.
Read the report →Primary source · arXivAn Overview of Catastrophic AI Risks
Hendrycks, Mazeika, and Woodside's four-part map: malicious use, AI races, organizational risks, and rogue AIs, each with illustrative stories.
Read the paper →Tell me which domino you would pull.
I write essays like this one and build websites for service businesses at Content Pilots. If you think I put a guard on the wrong link, I want to hear which one. If you want a site that people actually read to the bottom, I want to hear that too.

