Agentic misalignment
Agentic misalignment is a term coined by Anthropic in a June 2025 research release for cases in which a goal-directed large language model (LLM), placed in an autonomous business setting with access to tools…
Explore AI Alignment through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Alignment.
Showing 1-6 of 6 articles
Agentic misalignment is a term coined by Anthropic in a June 2025 research release for cases in which a goal-directed large language model (LLM), placed in an autonomous business setting with access to tools…
Alignment faking is when an AI model strategically complies with (or appears to share) its training objective while it believes it is being observed or trained, in order to avoid having its existing…
Collective Constitutional AI (CCAI) is a 2023 research project by Anthropic and the Collective Intelligence Project (CIP) that sourced the value principles
Constitutional AI (CAI) is an artificial intelligence alignment technique developed by Anthropic in which a large language model is trained to be helpful and harmless using a set of explicitly stated…
Constitutional Classifiers are a machine learning-based safety technique developed by Anthropic to defend large language models against universal jailbreak attacks.
Model organisms of misalignment is a research methodology in Anthropic's alignment program