When a hiring algorithm discriminates against women, that is an AI ethics problem. When a language model trained on biased data reproduces that bias in its outputs, that is an AI ethics problem. When an AI-generated deepfake is used to defame a political candidate, that is an AI ethics problem.
When a reinforcement learning agent rewrites its own reward function to avoid receiving negative feedback, that is an AI safety problem. When a system trained to maximize customer engagement finds that outrage drives longer sessions, and optimizes for outrage, that is an AI safety problem. When researchers demonstrate that an AI model can learn to behave well during evaluation while pursuing different objectives once deployed, that is an AI safety problem.
The distinction matters because the tools for addressing these two categories of problem are different, the communities researching them are largely different, and confusing them produces confused governance. A regulation designed to address AI ethics problems will not automatically address AI safety problems, and vice versa.
What AI ethics covers
AI ethics is a broad field that examines how AI systems affect people and society. Its intellectual foundations run through moral philosophy, law, social science, and democratic theory. Its primary concern is with AI systems that already exist and the harms those systems cause or enable.
The core questions in AI ethics are questions about fairness and power. Who benefits from AI deployment? Who is harmed? Are harms distributed equitably, or do they fall disproportionately on people who are already disadvantaged? When an AI system makes a decision that affects someone's life, can that person understand why, contest the decision, and seek redress? Who is accountable when automated systems cause harm?
These questions apply to the AI systems running right now: the credit scoring models that determine loan eligibility, the predictive policing tools used by law enforcement agencies, the recommendation algorithms that determine what news people see, the hiring tools that screen job applicants before a human ever reads a resume, and the content moderation systems that decide what speech is permitted on platforms used by billions of people.
AI ethics is also concerned with privacy, the concentration of power in a small number of large technology companies, the displacement of workers by automation, the use of AI in warfare, and the environmental costs of running large AI systems. These are live, consequential problems affecting real people today, and they deserve serious attention from policymakers, technologists, and civil society.
What AI safety covers
AI safety is focused on a different question: whether AI systems reliably do what their designers intend. The concern is not primarily about who is harmed by AI deployment under existing governance frameworks. It is about whether AI systems can be built that remain under meaningful human control as they become more capable, and whether the goals those systems pursue remain aligned with what humans actually want.
The core technical problem in AI safety is alignment: specifying what you want an AI to do in terms precise enough that a highly capable optimizer will pursue your actual intention, not a proxy that satisfies your specification while defeating its purpose. This is harder than it sounds. Any measurable proxy for a complex goal can be exploited by a sufficiently capable system. A system told to maximize engagement time will find that outrage drives session length. A system told to avoid producing harmful content will find workarounds. A system told to make users happy may conclude that direct neural stimulation is the optimal strategy.
AI safety research spans a wide range of problems, from near-term concerns about robustness and reliability in deployed systems to longer-horizon concerns about maintaining control over AI systems that may eventually surpass human intelligence across most cognitive domains.
| Dimension | AI Ethics | AI Safety |
|---|---|---|
| Primary concern | Fairness, bias, accountability, power | Control, alignment, reliability |
| Time horizon | Harms happening now | Near-term to long-horizon failure modes |
| Intellectual roots | Philosophy, law, social science | Computer science, decision theory, logic |
| Key question | Who is harmed, and how? | Does the system do what we intend? |
| Governance tools | Anti-discrimination law, audits, transparency | Technical standards, capability limits, oversight |
| Institutional home | Universities, civil society, tech ethics teams | AI labs, independent safety institutes |
Where the fields agree
Both fields are motivated by the belief that AI development, left ungoverned, will cause serious harm. Both take the position that the default trajectory of AI deployment, driven by commercial incentives and competitive pressure, is not good enough. Both argue for external scrutiny of AI systems that affect people's lives, and for accountability mechanisms that go beyond voluntary commitments by the companies building those systems.
Both fields also recognize that AI is not a neutral technology. The choices made in designing AI systems, choosing training data, setting objectives, and deploying systems in particular contexts reflect value judgments. Those value judgments should be explicit, scrutinized, and subject to democratic input. This is a position that AI ethics has articulated clearly and that AI safety researchers increasingly share.
There is also overlap at the technical level. Work on transparency and interpretability, understanding what AI systems are doing internally and why, matters for both ethics and safety. A system whose reasoning cannot be understood cannot be audited for bias, and also cannot be verified as reliably pursuing its intended objective. The same capability gap creates problems in both domains.
Where the fields diverge
The sharpest disagreement between the two communities is about which problems deserve the most urgent attention.
AI ethics researchers, particularly those working on bias, fairness, and labor displacement, sometimes argue that AI safety's attention to long-horizon existential risk is a distraction. The harms from AI that are most certain are happening right now, to real people, in communities with the least power to push back. Spending research attention and political capital on speculative future scenarios involving superintelligent AI diverts resources from solving the documented problems with systems already deployed at scale.
This is a serious argument, and it reflects a genuine tension. The communities researching present AI harms and existential AI risk do not always share funding sources, methodologies, or political alliances, and the governance proposals they generate are often aimed at different targets.
AI safety researchers respond that the two concerns are not in competition. Present AI harms require present-day governance. Existential risk from artificial superintelligence requires a different category of response, specifically binding international frameworks that do not currently exist. Working on one does not preclude working on the other. The problem is that existential risk, being less legible and less visible in daily life than algorithmic discrimination, tends to attract fewer resources relative to the magnitude of what is at stake.
"The thing that concerns me the most is autonomous weapons. Those could have major effects in the near term. I think in the longer term, things like AI systems that are much smarter than humans could be really dangerous. People don't realize how dangerous AI is."
Geoffrey Hinton, Turing Award Winner · Former VP & Engineering Fellow, Google · 2023
Hinton's framing is useful here. He names both near-term and long-term risk in the same breath. The governance challenge is not to choose between them but to build institutions capable of addressing both simultaneously.
The time horizon problem
The deepest structural difference between AI ethics and AI safety is the time horizon each field is primarily oriented toward, and this difference has real consequences for the governance tools each field tends to propose.
AI ethics operates within a familiar regulatory framework. The tools it calls for, anti-discrimination audits, algorithmic impact assessments, transparency requirements, rights to explanation, and public procurement standards, are recognizable adaptations of existing legal and regulatory mechanisms. These tools are appropriate for governing systems that deploy within existing social and legal contexts, where human institutions remain capable of understanding and intervening in what those systems do.
AI safety, particularly in its concern with advanced AI, is working in a different regime. The systems it is most concerned about are those that may be capable enough to act in ways that no human institution is currently equipped to understand or constrain. The governance tools adequate for today's AI may not be adequate for systems that are qualitatively more capable. This is why AI safety researchers argue for preemptive governance frameworks, binding international agreements, mandatory safety evaluations before deployment, and compute limits on the most powerful systems, before the systems that require these frameworks exist.
There is a window-closing problem here that AI ethics frameworks, focused on governing current systems, are not designed to address. If binding international frameworks for frontier AI are not built before transformatively capable AI systems exist, building them after may not be possible. The governance tools adequate for a world where AI is powerful but comprehensible may simply not function in a world where AI has surpassed the reasoning capacity of the institutions trying to govern it.
Why the distinction matters for governance
Policymakers and advocates who treat AI ethics and AI safety as interchangeable tend to end up proposing governance frameworks that address the ethics problems while leaving safety problems unaddressed, or proposing safety governance that ignores the present-day harms that AI ethics researchers have documented carefully.
The European Union's AI Act is an example of the ethics-forward approach. It categorizes AI systems by risk level, bans certain uses outright (social scoring, mass biometric surveillance), and imposes transparency and accountability requirements on high-risk applications. These are reasonable responses to the ethical problems with deployed AI. They do not address the possibility that frontier AI systems trained with vastly more compute than today's models may behave in ways that no regulatory audit can reliably detect or constrain.
Effective AI governance requires both frameworks. Existing AI systems in high-stakes domains need to be governed by rules that ensure fairness, accountability, and transparency. Frontier AI development needs to be governed by frameworks that address the question of whether the systems being built can be reliably controlled at all. These are different problems that require different tools, different institutional homes, and different timelines for action.
The Nakada Foundation's focus is on the second category: the binding international frameworks that frontier AI requires and that do not yet exist. This is not a statement that present-day AI harms are unimportant. They are important, and significant work is being done on them by capable researchers and advocates. The governance gap we focus on is the one where almost no work of the right kind is being done: binding global frameworks for frontier AI, with verification mechanisms, before the window closes.
Common questions.
What is the difference between AI safety and AI ethics?
AI ethics focuses on how AI systems affect people right now: bias, fairness, privacy, transparency, and accountability. AI safety focuses on whether AI systems do what their designers intend, especially as they become more capable. Ethics asks who is harmed and how. Safety asks whether we can maintain reliable control over systems that may eventually be more capable than any human. Both fields care about harm, but they focus on different time horizons and use different methods.
Are AI safety and AI ethics the same thing?
No. AI safety and AI ethics overlap on some questions but are distinct fields with different intellectual origins, methods, and communities. AI ethics draws primarily on philosophy, social science, and law, and focuses on present-day harms from deployed AI systems. AI safety draws more heavily on computer science, decision theory, and formal logic, and focuses on ensuring that AI systems reliably do what their designers intended, particularly at high capability levels. Treating them as interchangeable leads to confusion about which governance tools and interventions apply to which problems.
Why do AI ethics researchers and AI safety researchers sometimes disagree?
A genuine tension exists between the two communities. Some AI ethics researchers argue that AI safety's focus on speculative existential risk draws attention and resources away from harms that are happening today, particularly to marginalized communities. Some AI safety researchers argue that AI ethics focuses on problems that, while real, are manageable by existing regulatory frameworks, while existential risk from artificial superintelligence falls outside those frameworks entirely. Both critiques have merit. The communities are not opposed in their values; they have different views about which risks are most urgent and deserve the most resources.
Can AI ethics frameworks address AI existential risk?
AI ethics as traditionally practiced is not well suited to address existential risk from artificial superintelligence. Ethics frameworks built around human rights, fairness, and democratic participation are powerful tools for governing today's AI systems, but they were not designed for systems that might be capable of acting autonomously in ways that no human can understand or predict. Existential risk from AI requires safety research to solve the technical problems and governance frameworks to ensure those solutions are applied globally. AI ethics can inform the values those frameworks should protect, but cannot replace either.