AI Safety Levels
AI Safety Levels (ASL) are a tiered risk-classification scheme developed by Anthropic to grade frontier AI systems by their potential for catastrophic harm and to attach progressively stricter safety…
Explore Anthropic through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Anthropic.
Showing 1-16 of 16 articles
AI Safety Levels (ASL) are a tiered risk-classification scheme developed by Anthropic to grade frontier AI systems by their potential for catastrophic harm and to attach progressively stricter safety…
Agentic misalignment is a term coined by Anthropic in a June 2025 research release for cases in which a goal-directed large language model (LLM), placed in an autonomous business setting with access to tools…
Alignment faking is when an AI model strategically complies with (or appears to share) its training objective while it believes it is being observed or trained, in order to avoid having its existing…
The Anthropic Long-Term Benefit Trust (LTBT) is a Delaware purpose trust that holds a special class of Anthropic stock and uses it to elect a portion of the company's board of directors.
Claude Code Review is a multi-agent code review system developed by Anthropic that automatically analyzes GitHub pull requests for bugs, security vulnerabilities, and logic errors.
Constitutional AI (CAI) is an artificial intelligence alignment technique developed by Anthropic in which a large language model is trained to be helpful and harmless using a set of explicitly stated…
Constitutional Classifiers are a machine learning-based safety technique developed by Anthropic to defend large language models against universal jailbreak attacks.
Dario Amodei (born 1983) is an Italian-American artificial intelligence researcher, entrepreneur, and the co-founder and CEO of Anthropic
Evan Hubinger is an American AI safety researcher who leads the alignment stress-testing team at Anthropic, where he serves as a Member of Technical Staff and manager.
GTG-1002 is the internal tracking name that Anthropic gave to a cyber espionage operation it says it detected in September 2025 and disclosed in November 2025
Many-shot jailbreaking is a technique for bypassing the safety training of a large language model by filling its context window with a long series of faux dialogue turns in which an AI assistant complies with…
Model organisms of misalignment is a research methodology in Anthropic's alignment program
Petri is an open-source tool, released by Anthropic, that automates parts of AI alignment auditing by using one AI model to probe another and a third to score what happens.
Sabotage evaluations are a set of tests, introduced by Anthropic in October 2024, that try to measure whether a frontier language model is capable of covertly subverting human oversight, decision-making, and…
"Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training" is a January 2024 Anthropic research paper (arXiv:2401.05566, submitted 10 January 2024) that deliberately trained large language…
The Anthropic Institute is a research organization within Anthropic dedicated to studying the societal challenges posed by increasingly powerful artificial intelligence systems.