After a serious near-miss, an investigator from the National Transportation Safety Board does not ask the airline to file a private note and move on. The event enters a system. Findings are written so designers, trainers, and regulators can all change behavior. The passenger who never hears the story still flies inside the update. That shared memory is a large part of why commercial flying became as safe as it is.
AI incident reporting is the attempt to give this field the same habit: systematic recording, sharing, and investigation of cases where systems cause harm or come close. When a model fails in deployment, slips its guardrails under misuse, or shows a dangerous behavior in testing, there is still no consistent duty to record it, no shared place that must receive it, and no standard investigation path. Lessons that could protect everyone stay inside one company, or disappear.
The proposal is simple on paper: define categories of harm and near-miss, require reporting into a shared system, and let the field learn once, collectively.
What a good system would capture
The useful version is broader than headline disasters.
- Deployed harms, where a system in the real world caused damage or serious malfunction.
- Near-misses, where something went wrong and harm was narrowly avoided. Aviation treats these as informative as wrecks.
- Dangerous behaviors found in testing, including failures surfaced by red teaming and evaluations, so one lab's hazard warning reaches the rest.
- Security events, such as attempts to steal model weights or circumvent safeguards.
Reporting has to be structured and, where appropriate, protected. Organizations disclose more honestly when safety investigation is separated from blame theater. Aviation learned that split for a reason. Without it, the rational move is silence.
Why it is worth doing
Shared incident data lets the field spot patterns no single organization would see. Regulators get an evidence base grounded in what is actually going wrong rather than in speculation. One lab's failure mode becomes everyone's warning the week it is found. Over time the industry gains the institutional memory it still lacks. Among AI governance measures this one is cheap, widely liked, and overdue. The Foundation wants it mandatory.
The folk objection, named
The optimistic version says aviation proved the method: keep reporting, keep investigating, and catastrophic risk will fall the same way hull-loss rates fell. On this view, incident systems are the main safety engine, and anticipatory limits are a distraction until we have more crash data.
Notice the tense that engine runs on. Incident reporting is retrospective. It learns from harms that have already happened. That works when failures, however tragic, are survivable at the level of the industry, so each one can teach the fleet. The system improves event by event.
That logic breaks for the risks the Foundation is most concerned with. A catastrophic failure of a superintelligent system is not a board you convene afterward. The core claim of existential risk is that the most serious failure may be the one with no recovery and therefore no lesson. Incident reporting handles accumulating, survivable harms well. By design it cannot address the unrecoverable one.
Learning from failures assumes you survive them. For the failures that matter most, that assumption is the problem.
Where it sits
Build the system. Mandate it. Protect honest disclosure. Then keep it out of the top slot of the toolkit. It should sit beside the anticipatory limits and verification set out in 我们的计划, which exist precisely because some failures cannot be handled after the fact.
Common questions.
What is AI incident reporting?
AI incident reporting is the systematic recording, sharing, and investigation of cases where AI systems cause harm or nearly do. It is modeled on aviation, where every crash and near-miss triggers an investigation whose findings are published and fed back into the whole industry. The aim is a duty to report defined categories of AI harm and near-miss into a shared system, so the field can learn from failures once, collectively, rather than having each lesson stay locked inside a single company or disappear.
What should an AI incident reporting system capture?
More than obvious disasters. A good system records deployed harms where a real-world system caused damage, near-misses where harm was narrowly avoided, dangerous behaviors found in testing such as those surfaced by red teaming and evaluations, and security events like attempts to steal model weights or bypass safeguards. Reporting should be structured and, where appropriate, protected from blame, so organizations disclose honestly rather than hiding failures to avoid embarrassment or liability.
Why is AI incident reporting valuable?
Because shared incident data lets the field detect patterns no single organization would see, gives regulators an evidence base grounded in what is actually going wrong rather than speculation, warns everyone about a failure mode as soon as one party hits it, and builds an institutional memory the field currently lacks. It is one of the lower-cost, higher-consensus measures in AI governance, which is why making it mandatory and well-designed attracts broad agreement.
What can't AI incident reporting do?
It is inherently retrospective: it works by learning from harms that have already happened. That is what made it so effective in aviation, where failures are survivable at the industry level so each one can teach the fleet. But that logic breaks for catastrophic risks from artificial superintelligence, where the most serious failure may be one there is no recovering from and therefore no learning from. Incident reporting handles accumulating, survivable harms very well and, by design, cannot address the unrecoverable one, so it must sit alongside anticipatory limits rather than replace them.