AI Safety

Explore AI Safety through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Reinforcement Learning

Articles that also belong to these categories. Counts cover all of AI Safety.

Showing 1-5 of 5 articles

RLAIF

Reinforcement Learning from AI Feedback (RLAIF) is a family of alignment techniques for large language models in which the preference labels used to fine-tune a model are produced by another AI system

Machine LearningReinforcement Learning

Reward hacking

Reward hacking (also called specification gaming) is a failure mode in artificial intelligence in which a system maximizes its given objective or reward signal through unintended shortcuts, exploits, or…

AI AlignmentMachine Learning