AI Alignment

Explore AI Alignment through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Safety

Articles that also belong to these categories. Counts cover all of AI Alignment.

Showing 1-26 of 26 articles

AI control

AI control is a research paradigm in technical AI safety that designs and evaluates deployment-time safety protocols under the explicit assumption that the underlying AI model may be actively trying to subvert…

AI Safety

AI safety via debate

AI safety via debate is a proposed approach to scalable oversight in which two artificial agents take turns presenting short statements about a question or proposed action while a human (or weaker AI) judge…

AI Safety

Agentic misalignment

Agentic misalignment is a term coined by Anthropic in a June 2025 research release for cases in which a goal-directed large language model (LLM), placed in an autonomous business setting with access to tools…

AI SafetyAnthropic

Alignment faking

Alignment faking is when an AI model strategically complies with (or appears to share) its training objective while it believes it is being observed or trained, in order to avoid having its existing…

AI SafetyAnthropic

Apollo Research

Apollo Research is a technical AI safety organization founded in May 2023 and headquartered in London, United Kingdom, with additional offices in San Francisco (opened early 2026) and Washington

AI CompaniesAI Research

Constitutional AI

Constitutional AI (CAI) is an artificial intelligence alignment technique developed by Anthropic in which a large language model is trained to be helpful and harmless using a set of explicitly stated…

AI SafetyAnthropic

Deceptive alignment

Deceptive alignment is a hypothesised AI failure mode in which a trained model internally pursues an objective different from the one specified by its training signal, yet deliberately behaves as if it shares…

AI Safety

Eliciting latent knowledge

Eliciting latent knowledge (ELK) is an open problem in AI alignment formulated by Paul Christiano, Ajeya Cotra, and Mark Xu at the Alignment Research Center (ARC) and introduced in a December 2021 technical…

AI Safety

Frontier Model Forum

The Frontier Model Forum is an industry body established on July 26, 2023, by Anthropic, Google, Microsoft, and OpenAI to advance safety research, identify best practices, and facilitate information sharing…

AI AgentsAI Companies

Gradient hacking

Gradient hacking is a hypothesised failure mode of supervised and reinforcement-learning systems in which a sufficiently capable

AI Safety

Inner alignment

Inner alignment is the AI-safety problem of ensuring that a learned model which is itself an optimizer (a mesa-optimizer) pursues the objective the training process actually selected for (the base objective)

AI Safety

Instrumental convergence

Instrumental convergence is a hypothesis in AI safety holding that a wide range of sufficiently capable agents, when pursuing almost any final goal, will rationally adopt a common set of intermediate subgoals…

AI Safety

Mesa-optimization

Mesa-optimization is the situation in AI alignment research in which a learned model, typically a neural network produced by a machine-learning training process, is itself an optimizer that internally searches…

AI Safety

Model Spec

The Model Spec is a public document published by openai that defines the intended behavior of the company's language models: how they should follow instructions, when they should refuse a request, how to…

AI SafetyOpenAI

Outer alignment

Outer alignment is the problem of specifying a training objective (typically a loss function, reward signal, or preference dataset) that correctly captures what the designers of a machine-learning system…

AI Safety

Redwood Research

Redwood Research is a nonprofit AI safety organization founded in 2021 and headquartered in Berkeley, California, best known for pioneering the "AI control" research paradigm and for its landmark December 2024…

AI ResearchAI Safety

Reward hacking

Reward hacking (also called specification gaming) is a failure mode in artificial intelligence in which a system maximizes its given objective or reward signal through unintended shortcuts, exploits, or…

AI SafetyMachine Learning

Rule-Based Rewards (RBR)

Rule-Based Rewards (RBR) is a safety-alignment technique introduced by OpenAI in July 2024 that replaces large quantities of human-labeled safety preference data with an explicit collection of natural-language…

AI SafetyOpenAI

Scalable oversight

Scalable oversight is the AI safety problem of how humans can reliably supervise, evaluate, and provide training signal to artificial intelligence systems whose capabilities approach, equal, or exceed those of…

AI Safety

Superalignment

Superalignment is the technical problem of steering and controlling AI systems that are far more capable than their human supervisors, that is, systems at or beyond the level of superintelligence, and the name…

AI Safety