The governance literature on AI safety has a bias toward successful analogies. The Chemical Weapons Convention is cited frequently. The Montreal Protocol appears in most serious discussions. The IAEA's nuclear safeguards system is held up as a model for what an AI monitoring agency might look like. These are valuable references, but they produce an incomplete picture of what international governance can and cannot accomplish.

The failures are equally instructive, and in some respects more so. They reveal the specific design choices and political conditions that cause governance regimes to break down or fall short. For AI governance, which has not yet been designed, understanding these failure modes is more actionable than studying success stories, because the design choices that will determine success or failure are still open.

Failure mode one: the entry-into-force trap

The Comprehensive Test Ban Treaty was opened for signature in September 1996. By 2026, 187 countries have signed it and 178 have ratified it. It has not entered into force.

The reason is Article XIV, which requires that the treaty enter into force only after ratification by all 44 states that had nuclear power plants or research reactors when the treaty was negotiated. This list of 44 Annex 2 states was designed to ensure that the treaty covered the states most relevant to nuclear testing. It has instead become a permanent veto held by eight countries: China, Egypt, India, Iran, Israel, North Korea, Pakistan, and the United States. Until all eight ratify, the CTBT cannot enter into force regardless of how many other states have done so.

The CTBT has not been useless. The Comprehensive Nuclear-Test-Ban Treaty Organization operates a global monitoring system of seismic, hydroacoustic, infrasound, and radionuclide stations that can detect nuclear tests with high reliability. This monitoring has produced real intelligence about the nuclear programs of non-signatory states, including North Korea's six nuclear tests between 2006 and 2017. The CTBTO's monitoring would not exist without the treaty framework, even in its current provisional application. But the treaty's core obligation, banning nuclear tests, is not legally binding on the states most likely to conduct tests, because it has never entered into force.

The Annex 2 Trap

The CTBT's entry-into-force provision requiring ratification by all 44 named states was designed to ensure universal coverage of the most relevant actors. It instead created a structure where any single relevant state can block the entire regime indefinitely. An AI governance treaty with a similar provision, requiring ratification by the United States, China, Russia, the EU, and all major AI-developing countries before entry into force, would face the same structural vulnerability. The treaty should enter into force among willing parties at a defined threshold, not require universal coverage as a precondition.

For AI governance, the CTBT failure points to a specific design choice: entry-into-force provisions. A treaty that requires the participation of all major AI-developing states before it takes effect hands each of those states a unilateral blocking power over the entire regime. The US, China, the EU, the UK, and perhaps five or ten other states could be listed as required parties, and any one of them could prevent the treaty from entering into force indefinitely.

The more robust design allows the treaty to enter into force among a critical mass of willing parties, creating an operating governance regime, while leaving the door open for non-parties to join later. Non-parties face exclusion from the treaty's benefits (technology sharing, market access preferences, cooperative safety research) which creates incentives to join. The treaty's effectiveness among existing parties meanwhile creates a demonstrated governance model that reduces uncertainty about what joining would entail. The Mine Ban Treaty and the International Criminal Court both used this approach.

Failure mode two: a prohibition without verification

The Biological Weapons Convention was opened for signature in 1972 and entered into force in 1975. It prohibits the development, production, and stockpiling of biological weapons and requires states to destroy any existing stockpiles. It has 183 state parties. It has no verification mechanism, no inspectorate, and no formal body with authority to investigate alleged violations or make findings of non-compliance.

The Soviet Union ratified the BWC in 1975. In 1973, two years before ratification, the Soviet Council of Ministers had approved the creation of Biopreparat, a network of nominally civilian research institutes that constituted the world's largest offensive biological weapons program. At its peak in the 1980s, Biopreparat employed roughly 60,000 scientists and produced plague, smallpox, and anthrax weapons at industrial scale, along with genetically modified agents designed to be resistant to antibiotics and vaccines. This program operated in direct and comprehensive violation of the BWC throughout the 1970s and 1980s, until it was revealed by Soviet defectors after the USSR's collapse. The BWC regime had no way to detect it.

This is the most consequential failure of arms control verification in the post-war period. A prohibition that cannot be verified is a norm, not an accountability mechanism. It creates a public standard of acceptable behavior that has real effects on state behavior in many cases, while being structurally unable to detect systematic violation by states with the means and motive to violate covertly.

Negotiations to add a verification protocol to the BWC ran from 1995 to 2001. The United States rejected the proposed protocol in July 2001, arguing that its inspection procedures were both too intrusive on legitimate commercial biotechnology and too weak to reliably detect violations. The negotiations collapsed and have not been revived in any form that could produce a binding verification mechanism. The BWC remains in force as a norm without a verification body.

What the BWC failure means for AI governance

The analogy between biological weapons and AI is not perfect, but several features of the BWC failure are directly transferable to AI governance.

Dual-use is the central verification challenge. Biological weapons are produced using laboratory equipment, knowledge, and processes that are largely identical to legitimate pharmaceutical and public health research. Distinguishing offensive from defensive research, or civilian from weapons applications, requires highly trained inspectors with access to confidential commercial and technical information, and still produces genuine ambiguity in many cases. This dual-use character made a verification protocol for the BWC both technically demanding and politically contentious in ways that verification for chemical or nuclear weapons, which have clearer signatures, was not.

AI has the same dual-use character at a more fundamental level. A large language model trained for natural language generation is technically indistinguishable from one trained for persuasion at scale. A model trained for scientific research assistance is technically identical to one trained for autonomous action in high-stakes environments. The capabilities that make advanced AI commercially valuable are the same capabilities that make it potentially dangerous. This means that AI governance verification faces the same challenge that defeated BWC verification: it must distinguish dangerous from legitimate applications of technically identical systems.

"The Biological Weapons Convention shows exactly what a prohibition without verification produces over time. A widely accepted norm, demonstrably violated at scale by a major power for twenty years, with no mechanism to detect or respond to the violation. For AI governance, this is not a theoretical concern. It is the default outcome if verification is deferred to a future protocol that may never arrive."

The lesson is not that verification is impossible for AI, but that verification must be built into the treaty from the outset rather than deferred to a later protocol. The BWC's verification negotiations failed after twenty-six years of effort. An AI governance treaty that establishes the prohibition first and defers the verification architecture risks the same outcome: a norm without an accountability mechanism, operating alongside development programs it cannot see.

Failure mode three: withdrawal without cost

North Korea acceded to the Nuclear Non-Proliferation Treaty in 1985, announced withdrawal in 1993 and again in 2003, and conducted its first nuclear test in 2006. The NPT's withdrawal provision, Article X, allows any state to withdraw on ninety days' notice if it decides that extraordinary events have jeopardized its supreme national interests. North Korea's withdrawal came after the IAEA found it in non-compliance with its safeguards obligations, and the withdrawal effectively terminated the IAEA's ability to monitor North Korea's nuclear program.

For AI governance, the NPT/North Korea precedent points to the importance of withdrawal cost. A treaty that can be exited on ninety days' notice with no penalty provides weak assurance, because a state that wants to pursue activities prohibited by the treaty can simply leave before those activities become visible. Withdrawal provisions in an AI governance treaty should include: a defined notice period long enough for treaty bodies and other parties to respond; retention of inspection rights for a defined period after withdrawal notification; and automatic suspension of benefits upon notification of intent to withdraw rather than upon effective withdrawal.

These design choices would not prevent withdrawal, but they would ensure that withdrawal is costly and visible rather than routine and silent. The goal is not to trap states in treaty commitments they oppose, but to ensure that the act of leaving the regime is a politically significant event that generates international attention and diplomatic response.

AI governance designers have the advantage of being able to study these failures before writing the relevant provisions. The CTBT, BWC, and NPT failures are all documented, analyzed, and available. The design choices that produced those failures are avoidable. The question is whether the political process that produces an AI governance treaty will allow the technical design insights from these failures to survive contact with the negotiating positions of the major powers.

Common questions.

Why has the Comprehensive Test Ban Treaty never entered into force?

The CTBT's entry-into-force provision requires ratification by all 44 states that had nuclear power plants or research reactors when the treaty was negotiated. Eight of those 44 states have not ratified: China, Egypt, India, Iran, Israel, North Korea, Pakistan, and the United States. Until all eight ratify, the treaty cannot enter into force regardless of how many other countries have done so. This gives each holdout state a unilateral power to block the entire regime indefinitely, which is the direct consequence of designing entry-into-force provisions that require universal participation among the most relevant actors.

Does the Biological Weapons Convention prohibit anything if it has no verification?

The BWC creates a real and widely accepted norm against biological weapons that affects state behavior in most cases. But without verification, it cannot detect systematic violation. The Soviet Union ran the world's largest offensive biological weapons program throughout the 1970s and 1980s in direct violation of the BWC, entirely undetected by the treaty regime. The program was revealed only by defectors after the USSR's collapse. A prohibition without verification is a norm, not an accountability mechanism, and the distinction matters most for exactly the cases where state actors have strong motives to violate covertly.

What is the most important design lesson from treaty failures for AI governance?

Two lessons stand out. From the CTBT: entry-into-force provisions should allow the treaty to take effect among willing parties at a critical mass threshold, not require universal participation as a precondition. From the BWC: verification must be built into the treaty architecture from the beginning, not deferred to a future protocol. The BWC's verification negotiations ran for twenty-six years and ultimately collapsed. An AI governance treaty that establishes the prohibition without the verification body risks the same outcome: a norm operating alongside development programs it cannot see.

Is the Nuclear Non-Proliferation Treaty a success or a failure?

A partial success. The NPT has created a strong international norm against nuclear weapons acquisition and kept the number of nuclear-armed states far below what analysts predicted in the 1960s. But India, Pakistan, and Israel never joined; North Korea withdrew and tested nuclear weapons; and the nuclear-armed P5 states have not disarmed as the treaty requires. The IAEA safeguards system has detected most non-compliance attempts, but not all. The NPT is better evaluated as an imperfect but consequential instrument than as a straightforward success, and this nuanced reading is more useful for AI governance design than either a triumphalist or dismissive interpretation.