Jason Wei and coauthors at Google and elsewhere catalogued tasks where performance sat near chance until model scale crossed a threshold, then rose sharply. Multi-step arithmetic, specialized-domain questions, instructions in languages barely present in training, chains of reasoning steps: none of these was designed in as a feature. They showed up as models grew. The paper called them emergent abilities.

An emergent ability, in that usage, is a capability a smaller model does not have and a larger one does, where the transition looks less like a smooth ramp and more like a switch. Below some scale the model cannot do the task. Past that scale it can. The ability was not readable from the trend of smaller systems alone.

A real debate worth flagging

Emergence is not settled science. Some researchers argue that a share of reported jumps is partly a measurement artifact: a harsh all-or-nothing metric makes steady underlying progress look sudden, while a smoother metric shows gradual improvement. That critique lands on specific cases and deserves a fair hearing.

It does not dissolve the practical problem. Whether a capability arrives as a true discontinuity or as a steep curve noticed only after it crosses a threshold, the governance situation is the same. We learn a model can do something new after it can, not before. For a policymaker deciding whether the next training run is safe to green-light, the distinction between genuine emergence and a sharp curve is academic. Either way the surprise sits on the far side of the decision.

Why unpredictability is the core issue

We cannot reliably predict, before training a larger model, the full list of things it will be able to do. Scaling laws forecast broad performance measures such as loss reasonably well. They do not name which specific abilities will appear, or when. Capability is more predictable in aggregate than in particulars.

For most abilities that is merely interesting. For safety-relevant ones it is alarming. The same unpredictability that yields a surprise talent for translation can yield a surprise talent for the tasks covered in dangerous capability evaluations: cyber assist, bioweapon help, manipulation at scale. A capability no one anticipated is a capability no one tested for and no one built safeguards against.

We can forecast roughly how good the next model will be on average metrics. We cannot forecast everything it will be able to do. That gap is where the risk sits.

What it means for how we scale

Emergence turns each large jump in scale into a step into partial darkness. You build a system whose complete capability profile you learn only after it exists. By then, if a dangerous capability came with it, the system already has it. That is a different risk from a known hazard you can measure and mitigate in advance.

The Foundation's case follows directly. Frontier scaling should not run ahead of the ability to understand what is being created. The burden should fall on demonstrating that a large new system is safe before it is trained and deployed, not on discovering its dangers afterward. When abilities can arrive unannounced, caution before the jump is the only caution that helps. That framework is in our plan. The shortening timelines that raise the stakes are covered in the AGI timeline.

Common questions.

What are emergent abilities in large language models?

Emergent abilities are capabilities that are absent in smaller models and appear in larger ones, often with a sharp transition rather than a smooth ramp. Below a certain scale a model performs at chance on the task; past that scale it can do it. Reported examples include multi-step arithmetic, specialized question answering, following instructions in barely-represented languages, and multi-step reasoning, none of which were deliberately designed in.

Are emergent abilities real or a measurement artifact?

There is genuine debate. Some researchers argue that part of reported emergence is an artifact of harsh all-or-nothing metrics, which make steady underlying progress look like a sudden jump, and that smoother metrics reveal gradual improvement. This critique applies to specific cases. It does not remove the practical issue, because whether a capability arrives as a true discontinuity or a steep curve, we typically discover it only after a model has it.

Why do emergent abilities matter for AI safety?

Because we cannot reliably predict, before training a larger model, the full set of things it will be able to do. Scaling laws forecast broad performance well but not which specific abilities will appear or when. For safety-relevant abilities, such as helping with cyberattacks or bioweapons, that unpredictability means a dangerous capability can arrive unannounced, in a system we have already built, without having been anticipated, tested for, or guarded against.

How should emergence affect the way we scale AI?

It suggests that each large increase in scale is a step into partial darkness, since a model's complete capability profile is only learned after it exists. That argues for not letting frontier scaling outrun our ability to understand what we are building, and for placing the burden on demonstrating that a large new system is safe before it is trained, rather than discovering its hazards after deployment when it already possesses them.