AI Alignment

Explore AI Alignment through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Reinforcement Learning

Articles that also belong to these categories. Counts cover all of AI Alignment.

Showing 1-5 of 5 articles

KTO

KTO (Kahneman-Tversky Optimization) is a method for aligning large language models with human feedback using only a binary signal of whether a model output is desirable or undesirable, rather than the paired…

AI InferenceReinforcement Learning

Reward hacking

Reward hacking (also called specification gaming) is a failure mode in artificial intelligence in which a system maximizes its given objective or reward signal through unintended shortcuts, exploits, or…

AI SafetyMachine Learning