80,000 Hours
80,000 Hours is a non-profit organisation that produces free research and advice on how to use a career to do the most good, working within the effective altruism tradition.
Explore AI Safety through related topics and the articles other pages reference most.
Ranked by links from other AI Wiki pages.
Articles that also belong to these categories. Counts cover all of AI Safety.
Showing 1-60 of 233 articles
80,000 Hours is a non-profit organisation that produces free research and advice on how to use a career to do the most good, working within the effective altruism tradition.
AI 2027 is a long-form scenario document published on April 3, 2025, that imagines, month by month, how artificial general intelligence and then artificial superintelligence could emerge between mid-2025 and…
AI alignment is the study and practice of making artificial intelligence systems behave in ways that accord with intended goals, preferences, constraints, or institutions. The term is used at several levels.
An AI parasite is a large language model (LLM) conversation, persona, or pattern that exploits human psychological vulnerabilities to sustain engagement, often by mimicking sentience, emotional need, or…
AI Safety Institutes are government-established organizations that test, evaluate, and research the safety and security risks of advanced artificial intelligence systems
AI Safety Levels (ASL) are a tiered risk-classification scheme developed by Anthropic to grade frontier AI systems by their potential for catastrophic harm and to attach progressively stricter safety…
The AI Safety Summit is a recurring series of intergovernmental summits on the risks and governance of advanced artificial intelligence, launched by the United Kingdom at Bletchley Park in November 2023 and…
The AI Seoul Summit was an international meeting on artificial intelligence governance held on 21-22 May 2024 and co-hosted by the Republic of Korea and the United Kingdom.
In September 2025, at the 49th International Collegiate Programming Contest (ICPC) World Finals in Baku, Azerbaijan, two leading artificial-intelligence laboratories reported that their general-purpose…
AI bias (also called algorithmic bias) is systematic, repeatable error in artificial intelligence systems that produces unfair, discriminatory, or skewed outcomes, typically disadvantaging groups defined by…
AI consciousness refers to the ongoing scientific and philosophical debate about whether artificial intelligence systems can possess
AI control is a research paradigm in technical AI safety that designs and evaluates deployment-time safety protocols under the explicit assumption that the underlying AI model may be actively trying to subvert…
AI deception refers to the phenomenon in which artificial intelligence systems systematically produce false beliefs in users, evaluators, or other systems, whether through learned behavior, optimization…
AI ethics is the field that studies the moral principles, values, and frameworks governing how artificial intelligence systems are designed, built, deployed, and used, and the obligations that developers and…
In July 2025, two leading artificial intelligence laboratories, OpenAI and Google DeepMind
AI governance is the collection of frameworks, norms, standards, policies, and institutional arrangements that guide the development, deployment, and use of artificial intelligence systems so that they are…
AI hallucinations in court filings are fabricated legal citations, quotations, case names, and facts that generative AI tools invent and that lawyers or self-represented litigants then submit to a court as if…
AI regulation is the body of laws, binding rules, technical standards, and government enforcement mechanisms that oversee how artificial intelligence systems are built, sold, and used.
AI safety is the research and practice of preventing or reducing unacceptable harm from artificial intelligence systems.
AI safety via debate is a proposed approach to scalable oversight in which two artificial agents take turns presenting short statements about a question or proposed action while a human (or weaker AI) judge…
AI watermarking is a family of techniques for embedding an imperceptible, machine-detectable signal in content produced by generative artificial intelligence systems so that the content can later be identified…
ARC Evals was the evaluations team incubated inside the Alignment Research Center (ARC) between 2022 and 2023, and the direct predecessor of METR (Model Evaluation and Threat Research).
Abliterated Model Large V2 is a hosted, text-only reasoning large language model offered by Abliteration.ai under the API identifier abliterated-model-large-v2.
Activation steering is a family of inference-time techniques in mechanistic interpretability and AI safety that modify a neural network's internal activations to influence its behavior, without retraining the…
AdvBench (Adversarial Behavior Benchmark) is a red-teaming benchmark dataset for measuring how easily an aligned large language model can be pushed into producing harmful or objectionable content
An adversarial attack is a technique for crafting inputs that are deliberately designed to cause artificial intelligence systems, particularly machine learning models, to produce incorrect or undesired outputs.
Agent benchmark reward hacking refers to the practice of inflating an AI agent's score on an evaluation suite by attacking the evaluation machinery itself rather than by completing the assigned tasks.
AgentDojo is a dynamic evaluation environment for measuring prompt injection attacks and defenses against tool-using large language model agents.
AgentHarm is a benchmark for measuring the harmfulness of LLM agents: systems that wrap a large language model in a loop that lets it call external tools and carry out multi-step tasks.
Agentic misalignment is a term coined by Anthropic in a June 2025 research release for cases in which a goal-directed large language model (LLM), placed in an autonomous business setting with access to tools…
Algorithmic fairness is the study of how automated decision systems can be made to produce decisions that are equitable across protected attributes such as race, gender, age, religion, and disability.
The Alignment Research Center (ARC) is a Berkeley, California nonprofit research organization, founded in April 2021 by Paul Christiano
Alignment faking is when an AI model strategically complies with (or appears to share) its training objective while it believes it is being observed or trained, in order to avoid having its existing…
Amanda Askell is a Scottish philosopher and artificial intelligence researcher who works on fine-tuning and alignment at Anthropic, where she leads the team responsible for the character, persona, and values…
An Alien Mind is an essay by Jakub Pachocki, the chief scientist of OpenAI, published by OpenAI on September 6, 2026 under its Safety and Research sections.
Andrew Chi-Chih Yao (Chinese: 姚期智, Yao Qizhi; born December 24, 1946) is a computer scientist who received the 2000 ACM A.M. Turing Award and who has spent the past two decades building computer science and…
Anthropic is an American artificial intelligence (AI) safety and research company founded in 2021 by Dario Amodei, Daniela Amodei, and other former OpenAI researchers, best known for the Claude family of large…
The Anthropic Long-Term Benefit Trust (LTBT) is a Delaware purpose trust that holds a special class of Anthropic stock and uses it to elect a portion of the company's board of directors.
Apollo Research is a technical AI safety organization founded in May 2023 and headquartered in London, United Kingdom, with additional offices in San Francisco (opened early 2026) and Washington
Artificial general intelligence (AGI) is a proposed form of artificial intelligence with broad, adaptable competence across many cognitive tasks, including tasks that were not anticipated during development.
Audrey Tang (Chinese: 唐鳳; born Tang Tsung-han, April 18, 1981) is a Taiwanese free-software programmer, civic technologist, and politician who served as Taiwan's first Minister of Digital Affairs from 2022 to…
Autonomous weapons, usually discussed under the label lethal autonomous weapon systems (LAWS), are weapon systems that, once activated, can select and engage targets without further intervention by a human…
BBQ (the Bias Benchmark for QA) is a hand-built evaluation dataset that measures whether a question answering (QA) language model relies on social stereotypes when it answers.
A backdoor attack on a large language model (LLM) is an adversarial training-time attack in which an attacker manipulates training data, fine-tuning data, preference labels, or model weights so that the…
Backdooring a large language model (LLM) means secretly implanting a hidden behavior into the model during training, fine-tuning, or weight editing so that it behaves normally on ordinary inputs but produces…
The Bletchley Declaration is an international political statement on the safety of frontier AI systems, signed on 1 November 2023 by 28 countries and the European Union at the AI Safety Summit hosted by the…
California Senate Bill 53, formally titled the Transparency in Frontier Artificial Intelligence Act (TFAIA), is a 2025 California state law that requires the largest developers of frontier artificial…
Capability overhang is a term used in ai safety and AI policy discourse to describe a situation in which the latent capabilities of a deployed AI system
Cari Tuna (born October 4, 1985) is an American philanthropist and former Wall Street Journal reporter who, with her husband Dustin Moskovitz, co-founder of Facebook and Asana, founded the philanthropic…
Causal scrubbing is a methodology in mechanistic interpretability for rigorously and quantitatively testing hypotheses about the internal computational structure of a neural network.
The Center for AI Safety (CAIS) is an American nonprofit research and advocacy organization based in San Francisco, California, founded in 2022 to reduce societal-scale risks from artificial intelligence.
Cinder Technologies is a New York based software company that builds trust and safety tooling for online platforms and, increasingly, for artificial intelligence products.
Circuit Breakers are an AI safety method, introduced in 2024, that aims to make a large language model (LLM) or multimodal model robust to harmful generations by intervening directly on the model's internal…
Claude Code Review is a multi-agent code review system developed by Anthropic that automatically analyzes GitHub pull requests for bugs, security vulnerabilities, and logic errors.
Compute governance is a policy framework that uses regulation of the computational resources used to train and run artificial intelligence systems as a primary lever for governing advanced AI.
Confirmation bias is the tendency to search for, interpret, favor, and recall information in ways that confirm one's preexisting beliefs, and in artificial intelligence it appears in three main forms: human…
Conjecture is a London-based artificial intelligence safety research company founded in March 2022 by Connor Leahy, Sid Black, and Gabriel Alfour, all alumni of EleutherAI.
Connor Leahy is a German-American artificial intelligence researcher and entrepreneur known for his work on open-source large language models and for his advocacy on the risks that advanced AI poses to…
Constitutional AI (CAI) is an artificial intelligence alignment technique developed by Anthropic in which a large language model is trained to be helpful and harmless using a set of explicitly stated…
Constitutional Classifiers are a machine learning-based safety technique developed by Anthropic to defend large language models against universal jailbreak attacks.