Constitutional Classifiers
Constitutional Classifiers are a machine learning-based safety technique developed by Anthropic to defend large language models against universal jailbreak attacks.
Explore AI Alignment through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Alignment.
Showing 1-2 of 2 articles
Constitutional Classifiers are a machine learning-based safety technique developed by Anthropic to defend large language models against universal jailbreak attacks.
KTO (Kahneman-Tversky Optimization) is a method for aligning large language models with human feedback using only a binary signal of whether a model output is desirable or undesirable, rather than the paired…