AI Safety

Explore AI Safety through related topics and the articles other pages reference most.

Most referenced in this topic

Ranked by links from other AI Wiki pages.

Explore articles

Browse subtopics (47)

Articles that also belong to these categories. Counts cover all of AI Safety.

Showing 1-60 of 233 articles

80,000 Hours

80,000 Hours is a non-profit organisation that produces free research and advice on how to use a career to do the most good, working within the effective altruism tradition.

AI 2027

AI 2027 is a long-form scenario document published on April 3, 2025, that imagines, month by month, how artificial general intelligence and then artificial superintelligence could emerge between mid-2025 and…

AI Alignment

AI alignment is the study and practice of making artificial intelligence systems behave in ways that accord with intended goals, preferences, constraints, or institutions. The term is used at several levels.

AI EthicsMachine Learning

AI Parasite

An AI parasite is a large language model (LLM) conversation, persona, or pattern that exploits human psychological vulnerabilities to sustain engagement, often by mimicking sentience, emotional need, or…

Artificial Intelligence

AI Safety Levels

AI Safety Levels (ASL) are a tiered risk-classification scheme developed by Anthropic to grade frontier AI systems by their potential for catastrophic harm and to attach progressively stricter safety…

Anthropic

AI Safety Summit

The AI Safety Summit is a recurring series of intergovernmental summits on the risks and governance of advanced artificial intelligence, launched by the United Kingdom at Bletchley Park in November 2023 and…

Artificial Intelligence

AI at the 2025 ICPC World Finals

In September 2025, at the 49th International Collegiate Programming Contest (ICPC) World Finals in Baku, Azerbaijan, two leading artificial-intelligence laboratories reported that their general-purpose…

AI Companies

AI bias

AI bias (also called algorithmic bias) is systematic, repeatable error in artificial intelligence systems that produces unfair, discriminatory, or skewed outcomes, typically disadvantaging groups defined by…

AI EthicsArtificial Intelligence

AI control

AI control is a research paradigm in technical AI safety that designs and evaluates deployment-time safety protocols under the explicit assumption that the underlying AI model may be actively trying to subvert…

AI Alignment

AI deception

AI deception refers to the phenomenon in which artificial intelligence systems systematically produce false beliefs in users, evaluators, or other systems, whether through learned behavior, optimization…

Artificial Intelligence

AI ethics

AI ethics is the field that studies the moral principles, values, and frameworks governing how artificial intelligence systems are designed, built, deployed, and used, and the obligations that developers and…

AI EthicsArtificial Intelligence

AI regulation

AI regulation is the body of laws, binding rules, technical standards, and government enforcement mechanisms that oversee how artificial intelligence systems are built, sold, and used.

AI EthicsArtificial Intelligence

AI safety via debate

AI safety via debate is a proposed approach to scalable oversight in which two artificial agents take turns presenting short statements about a question or proposed action while a human (or weaker AI) judge…

AI Alignment

AI watermarking

AI watermarking is a family of techniques for embedding an imperceptible, machine-detectable signal in content produced by generative artificial intelligence systems so that the content can later be identified…

AI Policy & RegulationGenerative AI

AdvBench

AdvBench (Adversarial Behavior Benchmark) is a red-teaming benchmark dataset for measuring how easily an aligned large language model can be pushed into producing harmful or objectionable content

AI BenchmarksLarge Language Models

Adversarial attack

An adversarial attack is a technique for crafting inputs that are deliberately designed to cause artificial intelligence systems, particularly machine learning models, to produce incorrect or undesired outputs.

AgentDojo

AgentDojo is a dynamic evaluation environment for measuring prompt injection attacks and defenses against tool-using large language model agents.

AI AgentsAI Benchmarks

AgentHarm

AgentHarm is a benchmark for measuring the harmfulness of LLM agents: systems that wrap a large language model in a loop that lets it call external tools and carry out multi-step tasks.

AI AgentsAI Benchmarks

Agentic misalignment

Agentic misalignment is a term coined by Anthropic in a June 2025 research release for cases in which a goal-directed large language model (LLM), placed in an autonomous business setting with access to tools…

AI AlignmentAnthropic

Algorithmic fairness

Algorithmic fairness is the study of how automated decision systems can be made to produce decisions that are equitable across protected attributes such as race, gender, age, religion, and disability.

AI EthicsMachine Learning

Alignment faking

Alignment faking is when an AI model strategically complies with (or appears to share) its training objective while it believes it is being observed or trained, in order to avoid having its existing…

AI AlignmentAnthropic

Amanda Askell

Amanda Askell is a Scottish philosopher and artificial intelligence researcher who works on fine-tuning and alignment at Anthropic, where she leads the team responsible for the character, persona, and values…

AI EthicsPeople

An Alien Mind

An Alien Mind is an essay by Jakub Pachocki, the chief scientist of OpenAI, published by OpenAI on September 6, 2026 under its Safety and Research sections.

AI ResearchOpenAI

Andrew Yao

Andrew Chi-Chih Yao (Chinese: 姚期智, Yao Qizhi; born December 24, 1946) is a computer scientist who received the 2000 ACM A.M. Turing Award and who has spent the past two decades building computer science and…

Chinese AIComputer Science

Anthropic

Anthropic is an American artificial intelligence (AI) safety and research company founded in 2021 by Dario Amodei, Daniela Amodei, and other former OpenAI researchers, best known for the Claude family of large…

AI CompaniesArtificial Intelligence

Apollo Research

Apollo Research is a technical AI safety organization founded in May 2023 and headquartered in London, United Kingdom, with additional offices in San Francisco (opened early 2026) and Washington

AI AlignmentAI Companies

Audrey Tang

Audrey Tang (Chinese: 唐鳳; born Tang Tsung-han, April 18, 1981) is a Taiwanese free-software programmer, civic technologist, and politician who served as Taiwan's first Minister of Digital Affairs from 2022 to…

People

Autonomous weapons

Autonomous weapons, usually discussed under the label lethal autonomous weapon systems (LAWS), are weapon systems that, once activated, can select and engage targets without further intervention by a human…

AI EthicsAI Policy & Regulation

Backdooring LLMs

Backdooring a large language model (LLM) means secretly implanting a hidden behavior into the model during training, fine-tuning, or weight editing so that it behaves normally on ordinary inputs but produces…

Artificial Intelligence

Bletchley Declaration

The Bletchley Declaration is an international political statement on the safety of frontier AI systems, signed on 1 November 2023 by 28 countries and the European Union at the AI Safety Summit hosted by the…

AI Policy & Regulation

California Senate Bill 53

California Senate Bill 53, formally titled the Transparency in Frontier Artificial Intelligence Act (TFAIA), is a 2025 California state law that requires the largest developers of frontier artificial…

AI Policy & Regulation

Cari Tuna

Cari Tuna (born October 4, 1985) is an American philanthropist and former Wall Street Journal reporter who, with her husband Dustin Moskovitz, co-founder of Facebook and Asana, founded the philanthropic…

People

Causal scrubbing

Causal scrubbing is a methodology in mechanistic interpretability for rigorously and quantitatively testing hypotheses about the internal computational structure of a neural network.

Machine Learning

Claude Code Review

Claude Code Review is a multi-agent code review system developed by Anthropic that automatically analyzes GitHub pull requests for bugs, security vulnerabilities, and logic errors.

AI Code GenerationAnthropic

Compute governance

Compute governance is a policy framework that uses regulation of the computational resources used to train and run artificial intelligence systems as a primary lever for governing advanced AI.

AI Policy & Regulation

Confirmation Bias

Confirmation bias is the tendency to search for, interpret, favor, and recall information in ways that confirm one's preexisting beliefs, and in artificial intelligence it appears in three main forms: human…

AI EthicsData Science

Connor Leahy

Connor Leahy is a German-American artificial intelligence researcher and entrepreneur known for his work on open-source large language models and for his advocacy on the risks that advanced AI poses to…

People

Constitutional AI

Constitutional AI (CAI) is an artificial intelligence alignment technique developed by Anthropic in which a large language model is trained to be helpful and harmless using a set of explicitly stated…

AI AlignmentAnthropic