Anthropic published its first Responsible Scaling Policy in September 2023. Open the document and the structure shows up in the headings before the rhetoric does. Capability levels are named in advance. Each level is paired with safeguards that are supposed to switch on when evaluations say the model has crossed the line: harder weight security, tighter deployment rules, deeper red-teaming, and at the upper tiers a written commitment not to train or release until specified protections are in place. The industry short name is RSP. The underlying logic is often called if-then commitments.

The leading frontier labs published the best-known versions, and the format spread. Tiers typically run from systems roughly like today's through models that could meaningfully assist with a bioweapon, a serious cyberattack, or dangerous autonomy. The document is meant to answer a simple question before the emergency: when capability climbs, what exactly does the company do?

Why the format is a real improvement

Compared with a general promise to take safety seriously, an RSP is more concrete. It names capabilities, thresholds, and responses ahead of time, which turns a mood into something outsiders can point at. It is anticipatory. The lab has to decide what a dangerous capability will cost it before the rush of discovery and the press cycle arrive together.

The format also ties words to tests through dangerous capability evaluations, so thresholds have an operational meaning rather than a slogan. Some policies go further and commit to halt development if safeguards cannot keep pace with capability. Putting that sentence in a public document counts as a real advance.

The folk objection, named

The steelman says: public, specific, pre-committed rules are how industries mature; demanding external law before any internal policy is cynicism that freezes progress. On this view, RSPs are the draft constitution of frontier safety, and the right move is to encourage more of them rather than dismiss the authors.

True, and the weakness is still structural. The same lab writes the policy, defines the thresholds, runs the evaluations, judges whether a line was crossed, decides whether its safeguards suffice, and retains the power to revise the text. When commercial pressure to ship collides with a commitment the lab authored and administers, no independent party holds the whistle.

History does not favor the optimistic reading. The record of voluntary corporate safety pledges under competitive pressure, traced in our piece on voluntary commitments, is a record of standards that soften when they start to bite. An if-then rule is only as strong as the then, and the then is enforced by the party that benefits from weakening it. A rule you set, measure, and can amend for yourself is a statement of intent. It is not a constraint.

Two further gaps sit under the document. RSPs depend on evaluations to notice a crossed threshold, so everything that limits evaluations (including sandbagging, unknown capabilities, and the difficulty of proving safety) limits the policy built on them. And an RSP binds only the lab that adopts it. A competitor that declines, or a state program outside this culture, is untouched. No single company can close that coordination problem alone.

How the Foundation reads them

Keep the good machinery. Reuse the threshold logic, the if-then structure, and the link to evaluations. Change who holds the pen and who enforces the halt. Move those triggers from voluntary policy into binding law, backed by an independent body and by compute governance, so a lab cannot quietly revise its own ceiling when the ceiling becomes inconvenient.

Treat the policies as drafts of enforceable rules, then finish the job outside the lab: binding law, independent enforcement, compute controls. That transition is what our plan is for.

Common questions.

What is a responsible scaling policy?

A responsible scaling policy, or RSP, is a framework an AI lab publishes to govern its own development using an if-then structure: it defines levels of dangerous capability and commits in advance to specific safeguards that activate when a model reaches each level. Safeguards range from stronger security and deployment restrictions to commitments not to train or release a model until certain protections are in place. The approach is also described as if-then commitments.

Why are responsible scaling policies considered an improvement?

Because they are more concrete and anticipatory than a general promise to take safety seriously. They name specific capabilities, thresholds, and responses in advance, which makes them something an organization can be held against, and they force a lab to decide how it will handle a dangerous capability before it has one. They also tie thresholds to dangerous capability evaluations, and some commit to pausing development if safeguards cannot keep up with capability.

What is the main weakness of responsible scaling policies?

They are self-regulation. The same lab writes the policy, sets the thresholds, runs the evaluations, judges whether a threshold was crossed, decides whether its safeguards suffice, and can revise the policy. When commercial pressure to ship conflicts with a self-authored commitment, no independent party has the authority to enforce it. The history of voluntary corporate safety commitments under competitive pressure suggests such standards tend to soften just when they would otherwise bite.

How could responsible scaling policies be made stronger?

By moving their best ideas, the capability thresholds, the if-then structure, and the link to evaluations, from voluntary policy into binding law enforced by an independent body, and by backing them with compute governance. That keeps the useful structure while changing who holds the pen and who enforces the halt, so that a lab cannot quietly revise its commitments when they become inconvenient, and so that the rules bind competitors and state programs rather than only the lab that volunteered.