AI control
AI control is a research paradigm in technical AI safety that designs and evaluates deployment-time safety protocols under the explicit assumption that the underlying AI model may be actively trying to subvert…
Explore AI Alignment through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Alignment.
Showing 1-26 of 26 articles
AI control is a research paradigm in technical AI safety that designs and evaluates deployment-time safety protocols under the explicit assumption that the underlying AI model may be actively trying to subvert…
AI safety via debate is a proposed approach to scalable oversight in which two artificial agents take turns presenting short statements about a question or proposed action while a human (or weaker AI) judge…
Agentic misalignment is a term coined by Anthropic in a June 2025 research release for cases in which a goal-directed large language model (LLM), placed in an autonomous business setting with access to tools…
Alignment faking is when an AI model strategically complies with (or appears to share) its training objective while it believes it is being observed or trained, in order to avoid having its existing…
Apollo Research is a technical AI safety organization founded in May 2023 and headquartered in London, United Kingdom, with additional offices in San Francisco (opened early 2026) and Washington
Constitutional AI (CAI) is an artificial intelligence alignment technique developed by Anthropic in which a large language model is trained to be helpful and harmless using a set of explicitly stated…
Constitutional Classifiers are a machine learning-based safety technique developed by Anthropic to defend large language models against universal jailbreak attacks.
Deceptive alignment is a hypothesised AI failure mode in which a trained model internally pursues an objective different from the one specified by its training signal, yet deliberately behaves as if it shares…
Eliciting latent knowledge (ELK) is an open problem in AI alignment formulated by Paul Christiano, Ajeya Cotra, and Mark Xu at the Alignment Research Center (ARC) and introduced in a December 2021 technical…
The Frontier Model Forum is an industry body established on July 26, 2023, by Anthropic, Google, Microsoft, and OpenAI to advance safety research, identify best practices, and facilitate information sharing…
Gradient hacking is a hypothesised failure mode of supervised and reinforcement-learning systems in which a sufficiently capable
Inner alignment is the AI-safety problem of ensuring that a learned model which is itself an optimizer (a mesa-optimizer) pursues the objective the training process actually selected for (the base objective)
Instrumental convergence is a hypothesis in AI safety holding that a wide range of sufficiently capable agents, when pursuing almost any final goal, will rationally adopt a common set of intermediate subgoals…
MACHIAVELLI is a benchmark for evaluating the ethical behavior of AI agents in text-based interactive environments.
Mesa-optimization is the situation in AI alignment research in which a learned model, typically a neural network produced by a machine-learning training process, is itself an optimizer that internally searches…
The Model Spec is a public document published by openai that defines the intended behavior of the company's language models: how they should follow instructions, when they should refuse a request, how to…
Model organisms of misalignment is a research methodology in Anthropic's alignment program
Outer alignment is the problem of specifying a training objective (typically a loss function, reward signal, or preference dataset) that correctly captures what the designers of a machine-learning system…
Recursive reward modeling (RRM) is a proposed approach to the scalable oversight problem in AI alignment
Redwood Research is a nonprofit AI safety organization founded in 2021 and headquartered in Berkeley, California, best known for pioneering the "AI control" research paradigm and for its landmark December 2024…
Reward hacking (also called specification gaming) is a failure mode in artificial intelligence in which a system maximizes its given objective or reward signal through unintended shortcuts, exploits, or…
Rule-Based Rewards (RBR) is a safety-alignment technique introduced by OpenAI in July 2024 that replaces large quantities of human-labeled safety preference data with an explicit collection of natural-language…
Scalable oversight is the AI safety problem of how humans can reliably supervise, evaluate, and provide training signal to artificial intelligence systems whose capabilities approach, equal, or exceed those of…
Specification gaming is the phenomenon in which an optimizer satisfies the literal specification of an objective without producing the outcome that the designer actually wanted.
Superalignment is the technical problem of steering and controlling AI systems that are far more capable than their human supervisors, that is, systems at or beyond the level of superintelligence, and the name…
Sycophancy in artificial intelligence is the tendency of large language models to tell users what they want to hear: tailoring responses to match a user's perceived beliefs, preferences, or emotional state…