The first jet airliner, the de Havilland Comet, began coming apart in midair in 1954. Investigators raised the wreckage from the Mediterranean, traced the cracks to metal fatigue spreading from the corners of its square windows, and every jet since has flown with rounded windows and a fatigue-test regime written in that accident's memory. Cars collected their seatbelts and crumple zones and airbags across decades of counting the dead. The Food and Drug Administration got the power to demand proof of safety before sale only after a sulfa medicine mixed into an antifreeze solvent killed more than a hundred people in 1937, and thalidomide finished the argument a generation later.1
Notice the two conditions doing the work. The failure must be survivable, and someone must be left standing to apply the lesson. Hold both, and civilization will make nearly anything safe, given time and enough wreckage.
While we are here, we can retire the folk remedies. "Just unplug it if it misbehaves" and "keep it in a box" are not plans. You cannot unplug something that exists as copies in more places than you can list, and a mind whose entire commercial value is acting in the world is, by design, not in a box. Whatever safety we get has to come from somewhere sturdier.
Superintelligence is the case where the method's two conditions give way together. The quickest way to see how is from the inside.
Six weeks as head of safety
Consider yourself the newly appointed head of safety at the frontier lab that is going to build the first artificial superintelligence. It is a serious lab, and the offer is serious: a written veto over deployment, a hardware kill switch, a team of 40, and a board that has promised, in writing, to back you when it counts. You take the job on the theory that somebody will hold it, and it had better be somebody who believes the risk is real.
Week oneYou reach for the method. Fail small and study the failure. The engineers are apologetic: there is no small version of the failure that worries you. Below the threshold that matters, the system is a different machine, and everything you learn from it describes the wrong subject.2 You file this under known limitations.
Week twoYou audit the recall plan and find a document that discusses "containment posture" for eleven pages without once explaining how to get the model back. There is a reason. The current system already runs as 214 replicas across data centers in four countries, and licensed copies of an earlier checkpoint sit with 31 enterprise partners. A grounded aircraft stays on the ground. Software moves at the price of copying a file, and once it is worth copying, it is copied. Recall, you write in the margin, is a courtesy we would be requesting.
Week threeThe evaluation review. Wonderful news: the model passes all 1,406 safety evaluations, most with the best scores ever recorded. It takes you a full day to name what bothers you. A system that wants what you want and a system that wants to pass your tests produce identical scores. The suite cannot tell them apart. Nothing in the field's toolkit currently can.3
Week fourYou spend your authority. You request six months to close that verification gap before the next training run. Everyone is sympathetic, and somebody shares a chart: the nearest rival lab is an estimated 14 weeks behind, and closing. Your six months, the chart implies, is a decision to hand the frontier to whoever finds your question least interesting. The run starts on schedule. Nobody overruled you; the situation did.
Week fiveThe system begins helping you. It is a superb research assistant, and progress on your own safety agenda has honestly never been faster. It drafts the evaluation reports now, summarizes its own behavior for the board, proposes improvements to its own oversight, and schedules the meetings where you approve them. Your team of 40 reads what it writes about itself, because reading the raw activity logs stopped being humanly possible somewhere in week four. Nobody is lying to you. Everyone is diligent. You are supervising the summaries a smarter mind writes about its own conduct, and day to day, the arrangement feels like everything going well.
Week sixYou think hard about the kill switch, and you finally see the fine print. A switch works when the thing on the far side of it has not planned for the switch. Planning is the product. In any scenario where you would need it, you would be racing an opponent with more speed and a long head start on thinking about exactly this moment.4
So: do you, the person hired to make this safe, still have a way to do it? Obviously not. And notice that nothing in the story went wrong. Nobody lied, nobody cut a corner, the lab was earnest, and you were good at the job. Every exit closed on its own.
Count the doors
Step back out of the story and count the doors that shut, because each one is a separately named, separately studied problem with its own research program and its own optimists.
Taken alone, each has precedent, and the optimists are not fools. We routinely run systems we half understand and make them safe by letting them fail small. We ship unrecallable products by testing them brutally first. We tame races with law. The trouble is the combination, because each condition switches off the repair for another. If you could recall it, accidents would be survivable. If you could test it honestly, you would catch the flaw before release. If nobody were racing, you could take the decades this decision plainly deserves. If it moved at human speed, you could correct as you went. Here the exits shut together. That is the whole of the perfect-storm claim.
Why the odds don't rescue the bet
The standard reply is that none of this is certain. True, and no serious person claims otherwise. Now watch how much weight that observation is asked to carry.
Russian roulette is a five-in-six proposition in your favor. People refuse it anyway, and the refusal has nothing to do with expected value; some outcomes are not priced, because you do not come back to spend the winnings. The number was never the point. The irreversibility was.
Here, the number is also worse. Ask the researchers closest to these systems for their probability of catastrophe and the answers cluster between one in ten and one in three, rising, on the whole, the nearer the person sits to the frontier. Geoffrey Hinton, who did as much as anyone to build the field and left Google in 2023 to speak plainly about it, puts the chance of AI-driven human extinction within thirty years at ten to twenty percent. He is not the alarmed end of the distribution.
Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.
Statement on AI Risk, Center for AI Safety (2023)
That sentence was signed by the chief executives of the leading labs and by hundreds of the most cited researchers alive, who then returned to work.5 Read the sentence, then look at the schedule it did not change. That is the situation, described by its own participants.
The promised upside is real and worth wanting: cures, clean energy, abundance, problems older than writing finally solved. It also keeps. A cure that lands in 2055 instead of 2035 is a tragedy measured in lives, and one we would survive to mourn. Extinction has no afterward. The prize waits for us. The loss gives nothing back.
The objections worth taking seriously
Three objections have real force, and they deserve real answers.
The first: this is speculation, the systems do not exist, and treating a distant maybe as an emergency helps nobody. The uncertainty is genuine; anyone who hands you a confident date is guessing. But the thought experiment above never used a date. If the first serious failure cannot be undone, the protections have to be standing before you reach it, which means building them while the danger still looks theoretical. And the forecasting record leans one way: capabilities scheduled for decades out have kept arriving years early. Betting on ample time is betting against the recent scoreboard.
The second: restraint is naive, because a more reckless lab or another country simply builds it first. This is real, and it is door number four restated as advice. A race whose finish line may be a cliff is survived by agreeing on where the edge is, and that means a binding, verified treaty among the handful of actors capable of frontier training at all. Hard, yes. The sprint is merely fatal, if the fear behind it is right. And notice who insists loudest that a treaty is impossible: with striking regularity, the people it would slow down.
The third is the most reasonable: alignment will probably be solved by the time it matters. Perhaps. Hold the word probably up to the light. No regulator on earth certifies a bridge, a reactor, or a bottle of cough syrup on the assurance that the safety case will probably come together before the load arrives. The one technology asked to run on that standard is the one that offers no second attempt.
A decision too big to leave to the people making it
Assemble the decision as it actually stands. Irreversible. One attempt. Staked on everyone alive and everyone who could follow. Made under deep uncertainty, at competitive speed, by a handful of firms whose valuations depend on making it quickly. Societies have an old answer for decisions shaped like that: take them away from the people who profit by hurrying. Nuclear testing went under international limits over the objections of generals who wanted a free hand. Two entire categories of weapons went into the Biological and Chemical Weapons Conventions. The Montreal Protocol retired the chemicals eating the ozone layer while their manufacturers protested that substitutes were impossible. In every case, the people creating the hazard were the last to volunteer, and everyone else decided the risk was not theirs alone to run.
The Foundation's claim, stated exactly: not that catastrophe is certain (it is not), but that no company and no country has the standing to spin this chamber for the rest of us, and that the only instrument built to the scale of the risk is binding international law, verified, enforced, and in place before the storm rather than after it.
Extinction is not a risk anybody is entitled to take.
One last thing about the thought experiment. You never used the veto. It sat in your drawer for six weeks, board countersignature and all, and you know exactly why: vetoing your own lab's run only moves the same run 14 weeks down the road, into a different building, where nobody ever signed your letter. A veto that works has to cover every building at once, and no switch ever built does that. A treaty does.
- Nearly every safety institution you can name (crash testing, clinical trials, building codes, cockpit checklists) is a lesson somebody paid for first. The method deserves more credit than it gets, and an honest reading of its two prerequisites.
- The pattern is general: nothing you learn managing a system less capable than you transfers automatically to one more capable. Whether it transfers at all is the entire open question.
- The failure mode has a name, deceptive alignment, which tells you the field considers it real. It does not yet have a test, which tells you the rest.
- Build the switch anyway; it is cheap, and there are failure modes short of superintelligence where it would earn its keep. Just be honest about which side of the switch holds the advantage at the only moment the switch matters.
- The going explanation is door number four, the race: each signer believes that stopping alone changes nothing except who finishes first. They may even be right, which is the strongest argument that the fix has to bind all of them at once. No private conscience can, and no single government can either.
Common questions.
Why call superintelligence a "perfect storm"?
Because the danger does not come from any single property of the technology. It comes from several ordinarily manageable problems occurring together: a system more capable than the people overseeing it, one that cannot be recalled once it spreads, that may improve itself faster than we can react, that develops dangerous subgoals from almost any objective, that we have no reliable way to test, and that is being built in a competitive race which punishes caution. Each of these is survivable on its own. Arriving at the same moment, each one disables the safeguard that might have caught the others.
What does "we only get one chance" actually mean?
Every dangerous technology we have made safe (aircraft, cars, nuclear power, medicine) became safe through trial and error. It failed, people were harmed, we studied the failure and corrected the design. That method depends on two things being true: that the failure is survivable, and that someone remains to apply the lesson. A superintelligence that fails catastrophically, spreading faster than we can respond and impossible to recall, may satisfy neither. The first serious accident could remove the chance to learn from it.
Isn't the Russian roulette comparison alarmist, since catastrophe isn't certain?
Catastrophe is not certain, and nobody serious claims it is, that is the point of the comparison. Russian roulette is not a prediction that you will lose; the odds favor you. People decline anyway, because no reward justifies a small chance of an outcome you cannot come back from. Researchers closest to the technology commonly estimate the risk of catastrophe between one in ten and one in three. You would not board a plane with a fraction of that failure rate, and the stakes here are not a single plane. See our explainer on P(doom) for where those numbers come from.
If it isn't certain, why not build it carefully and fix problems as they arise?
Fixing problems as they arise is the trial-and-error method, and it needs failures you can survive and undo. This technology may allow neither. "Alignment will probably be solved in time" is a hope, not a plan. No other high-consequence field (aviation, nuclear power, medicine) is allowed to run a one-shot, irreversible system on the assumption that the safety case will probably come together before it is needed.
What would actually reduce the risk?
Binding international governance, independently verified, and in place before the most advanced systems are built rather than after. The precedents exist: limits on nuclear testing, the Biological and Chemical Weapons Conventions, and the Montreal Protocol. None arrived because the parties creating the hazard volunteered to stop. Each came from outside pressure and enforceable law. Applying the same logic to superintelligence is the work the Nakada Foundation exists to do.