Before a commercial airliner type is certified to carry passengers, regulators demand a structured safety argument: claims, reasoning, and evidence thick enough to reject. In the United Kingdom, nuclear site licensing still runs through a safety case the Office for Nuclear Regulation must accept before operations. Many medical devices follow the same burden. The operator does not get to run first and explain later. Permission follows a case that survived hostile review.

An AI safety case borrows that shape. It is a structured, evidence-backed argument that a particular system is safe enough to train or deploy in a particular setting. Applied to frontier models, it forces an affirmative case from the builder. Others no longer have to prove danger before anyone hits the brakes. The builder has to show safety, or stop.

What a real AI safety case would contain

A serious case is more than a checklist. It has a shape.

  • A claim: this model, used in this way, does not pose an unacceptable risk of specified harms.
  • An argument: the reasoning that connects evidence to the claim, including how identified risks are handled and why the safeguards are adequate.
  • Evidence: results from capability evaluations, red-teaming, security measures, and analysis the argument depends on.

It also has to state assumptions and failure points. A good case is falsifiable. It tells a reviewer what would need to be true for the conclusion to hold, so the reviewer can check whether those conditions actually hold. Without that, theater.

Why the discipline is valuable

Writing the argument down surfaces gaps a confident release note can hide. A risk is harder to wave away when you must construct explicit reasoning that addresses it. The method matches how other high-hazard fields got safer: by reasoning about hazards before operation, not only by counting wrecks afterward. It also slots into the threshold logic of responsible scaling policies, where a crossed capability line should trigger a case strong enough to justify the next step.

The folk objection, named

The pushback is practical: give the science a few more years of evaluations and interpretability, then safety cases will write themselves; requiring them now only slows useful systems. On this view, the paperwork should track the evidence, not lead it.

The uncomfortable reverse is the point. In aviation the case can lean on mature science, known failure rates, understood physics, and decades of fleet data. For a frontier model, an honest case runs into how little we can currently prove. We cannot yet demonstrate that a capable model is not deceptively aligned. We cannot rule out capabilities nobody thought to test. We cannot show that behavior in evaluation will hold in deployment, especially if the system is sandbagging.

A rigorous case for a sufficiently advanced model would need assurances current science cannot supply. So an honest attempt often produces, as its real output, a clear statement of why the system cannot yet be shown to be safe. The paperwork did its job when it refuses to rubber-stamp a gap.

The value of a safety case is the permission to refuse.

Why the Foundation supports them

A governance regime built on safety cases refuses to treat inability to prove danger as a license to proceed. Burden sits on the builder. If the builder cannot meet it, the answer is not to train the next jump and hope. Safety cases should be mandatory, independently reviewed rather than self-graded, and required before the largest training runs, not after a product launch.

Made binding in that form, they become one of the stronger tools available. That is why they feature in the wider design of our plan.

Common questions.

What is an AI safety case?

An AI safety case is a structured, evidence-backed argument that a specific AI system is safe enough to develop or deploy in a specific setting. The concept comes from safety-critical industries such as aviation, nuclear power, and medical devices, where operators must argue for safety before a regulator will let them run. Applied to AI, it requires a developer to make an affirmative case that a model is safe, rather than releasing it and waiting to see what happens.

What does an AI safety case include?

A serious safety case has three parts: a claim that the model, used in a defined way, does not pose an unacceptable risk of specified harms; an argument connecting evidence to that claim, including how risks are addressed and why safeguards suffice; and the supporting evidence, such as capability evaluations, red-teaming, and security measures. It should also state its assumptions and how it could fail, so a reviewer can check whether the conclusion actually holds.

Why are AI safety cases valuable?

Because the discipline of writing an explicit argument surfaces gaps that a confident announcement hides. It is much harder to wave away a risk when you must construct reasoning that addresses it. The approach also matches how other high-hazard fields became safe, by analyzing hazards before operating rather than testing after deployment, and it shifts the burden of proof onto the builder to show a system is safe rather than onto others to show it is dangerous.

Can we currently write a rigorous safety case for frontier AI?

Often not, and that is revealing. Unlike aviation, which can draw on mature science and decades of data, an honest safety case for a capable frontier model runs into how little we can prove: we cannot yet demonstrate a model is not deceptively aligned, cannot rule out untested capabilities, and cannot show that evaluation behavior will hold in deployment. A rigorous attempt therefore often produces, as its real output, a clear statement of why the system cannot yet be shown to be safe, which is itself a reason not to proceed.