AI Alignment

Explore AI Alignment through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Anthropic

Articles that also belong to these categories. Counts cover all of AI Alignment.

Showing 1-6 of 6 articles

Agentic misalignment

Agentic misalignment is a term coined by Anthropic in a June 2025 research release for cases in which a goal-directed large language model (LLM), placed in an autonomous business setting with access to tools…

AI SafetyAnthropic

Alignment faking

Alignment faking is when an AI model strategically complies with (or appears to share) its training objective while it believes it is being observed or trained, in order to avoid having its existing…

AI SafetyAnthropic

Constitutional AI

Constitutional AI (CAI) is an artificial intelligence alignment technique developed by Anthropic in which a large language model is trained to be helpful and harmless using a set of explicitly stated…

AI SafetyAnthropic