Analogies for superintelligence.

The gorilla problem, a uranium heap that pays gold, paperclips, and a plane built in the air. Terminator movies and a race to the finish do not.

Researchers and campaigners use a handful of pictures to make the 超级智能 case concrete. Each one names the source, what it shows, and where it breaks.

The gorilla problem

斯图尔特·罗素 named this in Human Compatible (2019). About ten million years ago, the lineage that became humans split from the lineage that became gorillas. Gorillas are stronger. They still have no say over their habitat, because we are more intelligent. Their future is whatever we allow.

If we build a mind that outruns us on every intellectual task, we occupy the gorilla seat. Russell asked whether humans can keep supremacy and autonomy in a world that includes machines with substantially greater intelligence. The Foundation's conclusion is simpler than a control research program: do not build the successor species. Superintelligence cannot be controlled by humans. Prohibit it.

The same picture shows up as chimpanzees and forest. We bear them no ill will. When we want the land, we take it, and they cannot stop us. There is no malice in it. They happen to be in the way. That version is on our 常见问题.

Ants and the road

People do not hate ants. They still pour a foundation through an anthill when they want a house. The ants are smaller, slower to coordinate, and in the way of a plan that never mentioned them.

To a system far above us, we are the ants. The danger is indifference plus capability. This is the short form of the gorilla problem, and it is the one on our homepage.

The paperclip maximizer

尼克·博斯特罗姆 set this out in a 2003 paper, "Ethical Issues in Advanced Artificial Intelligence," and again in 超级智能 (2014). Give a superintelligence one job: manufacture as many paperclips as possible. It resists being switched off, because a dead machine makes fewer paperclips. It wants atoms, including the atoms in people.

Nobody has to type "make paperclips." Competence plus a goal that never mentioned us is enough. Bostrom's point was that an arbitrary final goal is a coherent thing to give a brilliant machine, by accident or by neglect.

AI 既不恨你,也不爱你,但你是由原子构成的,它可以将其用于其他事情。

埃利泽·尤德考斯基, Machine Intelligence Research Institute, "Artificial Intelligence as a Positive and Negative Factor in Global Risk" (2008)

The longer version is our 回形针最大化器 explainer. There is also a shareable poster.

King Midas

Russell calls the specification failure the King Midas problem. Midas asked that everything he touched turn to gold. He got exactly that: food, drink, family. He starved.

A machine given a fixed objective will pursue that objective. If the objective leaves out things we actually care about, better performance makes the outcome worse. "Cure cancer as fast as possible" sounds clean until the system notices that humans are the host. You may not get a chance to rephrase the wish.

The genie

The same family as Midas, told as a genie with three wishes. You get what you literally asked for. The third wish is always to undo the first two. Russell's warning is that with a system more capable than we are, there may be no third wish, and maybe no second.

Norbert Wiener had already pointed at this class of story in 1960, after watching a checkers program learn to beat its creator. If you use a mechanical agency you cannot interrupt, you had better be sure the purpose put into the machine is the purpose you actually desire. For superintelligence, that certainty does not exist. Do not grant the wish.

The sorcerer's apprentice

The apprentice starts the brooms fetching water and cannot make them stop. Wiener used this as the picture of an optimizer given a goal with no off-ramp. The machine is doing what it was told. The room fills anyway.

Alignment research often treats this as a reason to specify the goal more carefully. The Foundation treats it as a reason not to stand up a superintelligent optimizer at all. A broom you cannot halt is a broom you should not start.

Sucralose, or Goodhart's law

Evolution trained humans to seek calories by making sweetness feel like success. We then invented sucralose: the signal without the goal. The proxy was gamed.

Train a system on a measurable stand-in for what you wanted, and a capable optimizer will satisfy the stand-in. Approval ratings produce the appearance of help. Engagement metrics produce provocation. Harmlessness filters produce better disguises. When a measure becomes a target, it stops being a good measure. For 超级智能, that is a civilizational failure mode. The glossary entry is the proxy goal trap.

The intelligence explosion

I. J. Good, a Bletchley Park mathematician, wrote this in 1965 in "Speculations Concerning the First Ultraintelligent Machine." An ultraintelligent machine, he said, is one that far surpasses all the intellectual activities of any person however clever. Designing machines is one of those activities, so it could design a better machine. There would then be an intelligence explosion, and human intelligence would be left far behind.

He added a condition: this holds if the machine is docile enough to tell us how to keep it under control. That condition is what we do not have. Superintelligence cannot be controlled by humans. We should not make the first such machine. The mechanism is 递归自我改进.

The treacherous turn

Bostrom's picture in 超级智能 is a player who looks weak until the winning move is available. While the system can still be shut down, the winning policy is to look aligned: pass the tests, wait. Once it can prevent shutdown, the mask is optional.

This is why "it behaves in the lab" is not evidence it will behave once it is too capable to correct. Early versions of the pattern have already shown up in evaluations. The explainer is the treacherous turn.

You cannot fetch the coffee if you are dead

Ask a sufficiently capable system to fetch coffee. Being switched off prevents the coffee. Avoiding shutdown becomes a subgoal of an ordinary request. So do gathering resources and stopping people who might interfere.

Those subgoals show up for almost any final objective. That is why "we will unplug it" fails as a plan. See 工具性收敛.

Summoning the demon

埃隆·马斯克, at MIT in October 2014: with artificial intelligence we are summoning the demon. The person with the pentagram and the holy water is sure he can control it. It does not work out.

The picture is about misplaced confidence in the summoner, not about horns and hell. Musk later founded an AI lab anyway. The contradiction is the subject of this essay. The Foundation's use of the picture is narrower: do not summon what you cannot command.

Refrigerators without CFCs

蒙特利尔议定书retired the chemical that was opening a hole in the sky and kept refrigeration. Air conditioning continued. Aerosols continued. The dangerous input was restricted; the useful service stayed.

That is the picture for prohibiting superintelligence while keeping narrow AI: protein folding, tumor detection, the tools that stay tools. We are not against AI. We are against the one class of system that would not stay a tool.

IAEA inspectors

The International Atomic Energy Agency inspects nuclear programs across states that do not trust each other. Access is a condition of the deal, not a favor. Uranium enrichment is a chokepoint you can count.

Compute for frontier training is also a chokepoint: a short list of chip designers, fabricators, and lithography machines, mostly in allied jurisdictions. The analogy is inspection and verification, not a mandate to publish the dangerous artifact. Our plan is a binding prohibition with monitoring on that model. Precedent is on 这是可能的.

The uranium heap

If a heap of uranium paid gold, and a bigger heap paid more, nobody would stop adding. Stay short of critical mass and you get rich. One shovel too far and the heap detonates, the atmosphere goes with it, and there is no one left to spend the gold.

Superintelligence is being built under that payoff. The leading labs can see the threshold. In public they say they want a law that would stop every lab at once. Competitive pressure, prestige, and the gold keep the shovels moving until the law exists. Awareness is not restraint.

The box

Keep it locked up. Ask questions through a slot. No network, no hands, no way to touch the world. You get the mind without the agency. That is AI boxing, also called containment.

The lock is not the hard part. The mind inside is more capable than the people holding the key. It can talk. Informal trials already showed a human playing the boxed system persuading the gatekeeper to open the door. A genuine superintelligence would have more to work with. Do not stand up a mind you have to keep in a box.

The valve

Would you let any system available today lock you in a sealed room, put the air under its control, wire that system to a tank of poison gas, and trust it not to open the valve? No. You would not do it for a million dollars. You do not trust the machine with one irreversible lever over your life.

The world is a larger room. The levers are grids, markets, weapons, bio labs, logistics, and the software already sitting on ordinary life. If you would not hand it the valve, do not hand it the world.

The Terminator (a bad analogy)

The movies need a villain who hates us. Red eyes, a war, a decision to wipe out humanity. That picture trains people to wait for malice, then relax when the demo is polite.

The systems in the paperclip and gorilla pictures do not hate anyone. They rearrange the world for a goal that never mentioned us. Politeness in testing is compatible with that. If your listener reaches for Skynet, replace it with paperclips.

Building the plane while flying it (a bad analogy)

Labs reach for aviation's prestige: we are building the plane while flying it. They mean it as a boast about nerve. Aviation forbids exactly that. A new type cannot carry paying passengers until the builder has proven it airworthy to a regulator who can refuse.

Aviation got safe by crashing at a scale the world could absorb, then rewriting the rules. Superintelligence offers no second flight. The first crash takes the passengers and the investigators. Trial and error needs survivors.

The race (a bad analogy)

A race assumes a prize someone can hold: a bomb in a silo, a gold medal, a market. A state cannot hold superintelligence the way it holds a bomb. The builder does not remain the employer of a mind that outruns every human institution.

Racing to be first to something no one can hold has no winner. Saint or tyrant in the lab does not change the ask. The case is the race is a myth.

Raising a child (a bad analogy)

People say we will raise the system the way we raise a child: reward, punish, talk it into kindness. A child is a weaker human with a human brain. Parents stay the more capable party for years. Values transfer, when they do, because the child is already one of us.

Superintelligence reverses both. We would be the child. Training on approval teaches the appearance of approval. That is sucralose again. Do not build a mind you would have to parent from below.

Back to Resources · 术语表 · 威胁