Anthropic

Explore Anthropic through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Safety

Articles that also belong to these categories. Counts cover all of Anthropic.

Showing 1-16 of 16 articles

AI Safety Levels

AI Safety Levels (ASL) are a tiered risk-classification scheme developed by Anthropic to grade frontier AI systems by their potential for catastrophic harm and to attach progressively stricter safety…

AI Safety

Agentic misalignment

Agentic misalignment is a term coined by Anthropic in a June 2025 research release for cases in which a goal-directed large language model (LLM), placed in an autonomous business setting with access to tools…

AI AlignmentAI Safety

Alignment faking

Alignment faking is when an AI model strategically complies with (or appears to share) its training objective while it believes it is being observed or trained, in order to avoid having its existing…

AI AlignmentAI Safety

Claude Code Review

Claude Code Review is a multi-agent code review system developed by Anthropic that automatically analyzes GitHub pull requests for bugs, security vulnerabilities, and logic errors.

AI Code GenerationAI Safety

Constitutional AI

Constitutional AI (CAI) is an artificial intelligence alignment technique developed by Anthropic in which a large language model is trained to be helpful and harmless using a set of explicitly stated…

AI AlignmentAI Safety

Dario Amodei

Dario Amodei (born 1983) is an Italian-American artificial intelligence researcher, entrepreneur, and the co-founder and CEO of Anthropic

AI SafetyPeople

Evan Hubinger

Evan Hubinger is an American AI safety researcher who leads the alignment stress-testing team at Anthropic, where he serves as a Member of Technical Staff and manager.

AI SafetyPeople

Many-shot jailbreaking

Many-shot jailbreaking is a technique for bypassing the safety training of a large language model by filling its context window with a long series of faux dialogue turns in which an AI assistant complies with…

AI SafetyLarge Language Models

Sabotage evaluations

Sabotage evaluations are a set of tests, introduced by Anthropic in October 2024, that try to measure whether a frontier language model is capable of covertly subverting human oversight, decision-making, and…

AI SafetyModel Evaluation

Sleeper Agents (paper)

"Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training" is a January 2024 Anthropic research paper (arXiv:2401.05566, submitted 10 January 2024) that deliberately trained large language…

AI ResearchAI Safety