AI Safety

Explore AI Safety through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Anthropic

Articles that also belong to these categories. Counts cover all of AI Safety.

Showing 1-16 of 16 articles

AI Safety Levels

AI Safety Levels (ASL) are a tiered risk-classification scheme developed by Anthropic to grade frontier AI systems by their potential for catastrophic harm and to attach progressively stricter safety…

Anthropic

Agentic misalignment

Agentic misalignment is a term coined by Anthropic in a June 2025 research release for cases in which a goal-directed large language model (LLM), placed in an autonomous business setting with access to tools…

AI AlignmentAnthropic

Alignment faking

Alignment faking is when an AI model strategically complies with (or appears to share) its training objective while it believes it is being observed or trained, in order to avoid having its existing…

AI AlignmentAnthropic

Claude Code Review

Claude Code Review is a multi-agent code review system developed by Anthropic that automatically analyzes GitHub pull requests for bugs, security vulnerabilities, and logic errors.

AI Code GenerationAnthropic

Constitutional AI

Constitutional AI (CAI) is an artificial intelligence alignment technique developed by Anthropic in which a large language model is trained to be helpful and harmless using a set of explicitly stated…

AI AlignmentAnthropic

Dario Amodei

Dario Amodei (born 1983) is an Italian-American artificial intelligence researcher, entrepreneur, and the co-founder and CEO of Anthropic

AnthropicPeople

Evan Hubinger

Evan Hubinger is an American AI safety researcher who leads the alignment stress-testing team at Anthropic, where he serves as a Member of Technical Staff and manager.

AnthropicPeople

Many-shot jailbreaking

Many-shot jailbreaking is a technique for bypassing the safety training of a large language model by filling its context window with a long series of faux dialogue turns in which an AI assistant complies with…

AnthropicLarge Language Models

Sabotage evaluations

Sabotage evaluations are a set of tests, introduced by Anthropic in October 2024, that try to measure whether a frontier language model is capable of covertly subverting human oversight, decision-making, and…

AnthropicModel Evaluation

Sleeper Agents (paper)

"Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training" is a January 2024 Anthropic research paper (arXiv:2401.05566, submitted 10 January 2024) that deliberately trained large language…

AI ResearchAnthropic