AI Alignment
AI alignment is the study and practice of making artificial intelligence systems behave in ways that accord with intended goals, preferences, constraints, or institutions. The term is used at several levels.
Explore AI Safety through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Safety.
Showing 1-26 of 26 articles
AI alignment is the study and practice of making artificial intelligence systems behave in ways that accord with intended goals, preferences, constraints, or institutions. The term is used at several levels.
AI bias (also called algorithmic bias) is systematic, repeatable error in artificial intelligence systems that produces unfair, discriminatory, or skewed outcomes, typically disadvantaging groups defined by…
AI consciousness refers to the ongoing scientific and philosophical debate about whether artificial intelligence systems can possess
AI ethics is the field that studies the moral principles, values, and frameworks governing how artificial intelligence systems are designed, built, deployed, and used, and the obligations that developers and…
AI regulation is the body of laws, binding rules, technical standards, and government enforcement mechanisms that oversee how artificial intelligence systems are built, sold, and used.
AI safety is the research and practice of preventing or reducing unacceptable harm from artificial intelligence systems.
Algorithmic fairness is the study of how automated decision systems can be made to produce decisions that are equitable across protected attributes such as race, gender, age, religion, and disability.
Amanda Askell is a Scottish philosopher and artificial intelligence researcher who works on fine-tuning and alignment at Anthropic, where she leads the team responsible for the character, persona, and values…
The Anthropic Long-Term Benefit Trust (LTBT) is a Delaware purpose trust that holds a special class of Anthropic stock and uses it to elect a portion of the company's board of directors.
Autonomous weapons, usually discussed under the label lethal autonomous weapon systems (LAWS), are weapon systems that, once activated, can select and engage targets without further intervention by a human…
BBQ (the Bias Benchmark for QA) is a hand-built evaluation dataset that measures whether a question answering (QA) language model relies on social stereotypes when it answers.
Confirmation bias is the tendency to search for, interpret, favor, and recall information in ways that confirm one's preexisting beliefs, and in artificial intelligence it appears in three main forms: human…
Effective altruism (often abbreviated EA) is a philosophical and social movement that uses evidence and careful reasoning to identify the most effective ways to benefit others
Existential risk from artificial intelligence (also called AI x-risk) is the hypothesis that the development of sufficiently advanced artificial intelligence could cause human extinction, permanent…
Human-in-the-loop (HITL) describes any arrangement in which a person is a required participant in an automated system's operating cycle rather than a bystander to it.
MACHIAVELLI is a benchmark for evaluating the ethical behavior of AI agents in text-based interactive environments.
Max Tegmark is a Swedish-American physicist and artificial intelligence researcher who is a professor of physics at the Massachusetts Institute of Technology (MIT) and the co-founder and president of the…
Meredith Whittaker is an American technologist, researcher, and privacy advocate who serves as president of the Signal Foundation, the nonprofit behind the encrypted messaging app Signal
Model welfare is the research area that investigates whether advanced AI systems might have morally relevant experiences or interests, such as suffering or wellbeing, and what (if anything) their developers…
Nick Bostrom (born Niklas Boström, 10 March 1973) is a Swedish-born philosopher best known for the 2014 book Superintelligence: Paths, Dangers, Strategies, the 2003 simulation argument, and the…
Responsible AI (RAI) is a framework for developing, deploying, and governing artificial intelligence systems in ways that are ethical, transparent, accountable, and aligned with human values.
Timnit Gebru is an Ethiopian-born computer scientist and a leading researcher in AI ethics, best known for co-authoring the 2018 "Gender Shades" study on bias in facial recognition, co-leading Google's Ethical…
Toby Ord is an Australian moral philosopher at the University of Oxford who founded the effective-altruism organisation Giving What We Can in 2009 and wrote the 2020 book The Precipice: Existential Risk and…
ToxiGen is a large-scale, machine-generated dataset designed for adversarial and implicit hate speech detection.
Transhumanism is an intellectual and cultural movement that holds that the human condition can and should be fundamentally improved through science and technology, in particular through technologies that…
William David MacAskill (born William Crouch; 24 March 1987) is a Scottish moral philosopher, author, and a co-founder of the effective altruism movement, best known as the leading public proponent of…