AI Safety

Explore AI Safety through related topics and the articles other pages reference most.

Explore articles

Browse subtopics (47)

Articles that also belong to these categories. Counts cover all of AI Safety.

Showing 181-233 of 233 articles

Purple Llama

Purple Llama is an umbrella project from Meta that bundles open, permissively licensed trust and safety tools and evaluations for generative AI.

Meta AIOpen Source AI

RLAIF

Reinforcement Learning from AI Feedback (RLAIF) is a family of alignment techniques for large language models in which the preference labels used to fine-tune a model are produced by another AI system

Machine LearningReinforcement Learning

Raine v. OpenAI

Raine v. OpenAI, Inc. is a wrongful-death lawsuit filed on August 26, 2025, in the Superior Court of California for the County of San Francisco (case number CGC-25-628528) by Matthew and Maria Raine, the…

AI Companies

Redwood Research

Redwood Research is a nonprofit AI safety organization founded in 2021 and headquartered in Berkeley, California, best known for pioneering the "AI control" research paradigm and for its landmark December 2024…

AI AlignmentAI Research

Reflection 70B controversy

The Reflection 70B controversy was an open-source AI credibility episode that began on September 5, 2024, when Matt Shumer, co-founder and chief executive of OthersideAI (the company behind the AI writing…

AI Companies

Remote Labor Index

The Remote Labor Index (RLI) is an AI benchmark that measures how well AI agents can complete real, paid remote knowledge work end to end.

AI Companies

Representation Engineering

Representation Engineering (often abbreviated RepE) is a top-down approach to artificial-intelligence transparency and control that reads and manipulates high-level concepts (such as honesty, harmlessness, and…

Interpretability

Responsible AI

Responsible AI (RAI) is a framework for developing, deploying, and governing artificial intelligence systems in ways that are ethical, transparent, accountable, and aligned with human values.

AI EthicsArtificial Intelligence

Responsible Scaling Policy

A Responsible Scaling Policy (RSP) is a self-imposed governance framework in which a frontier AI developer commits in advance to safety practices, capability evaluations, deployment restrictions, and security…

AI Policy & Regulation

Reward hacking

Reward hacking (also called specification gaming) is a failure mode in artificial intelligence in which a system maximizes its given objective or reward signal through unintended shortcuts, exploits, or…

AI AlignmentMachine Learning

Roman Yampolskiy

Roman Vladimirovich Yampolskiy (born August 13, 1979) is a Latvian-American computer scientist and a tenured associate professor in the Department of Computer Engineering and Computer Science at the University…

People

Rule-Based Rewards (RBR)

Rule-Based Rewards (RBR) is a safety-alignment technique introduced by OpenAI in July 2024 that replaces large quantities of human-labeled safety preference data with an explicit collection of natural-language…

AI AlignmentOpenAI

Rumman Chowdhury

Rumman Chowdhury (born April 1, 1980) is an American data scientist and social scientist known for her work in responsible AI, algorithmic auditing, and the study of algorithmic bias.

People

Sabotage evaluations

Sabotage evaluations are a set of tests, introduced by Anthropic in October 2024, that try to measure whether a frontier language model is capable of covertly subverting human oversight, decision-making, and…

AnthropicModel Evaluation

Safe Superintelligence Inc

Safe Superintelligence Inc. (abbreviated SSI) is an AI safety research company that is building a single product, a safe superintelligence, and nothing else.

AI Companies

Scalable oversight

Scalable oversight is the AI safety problem of how humans can reliably supervise, evaluate, and provide training signal to artificial intelligence systems whose capabilities approach, equal, or exceed those of…

AI Alignment

Seoul Declaration

The Seoul Declaration is the short name for the Seoul Declaration for Safe, Innovative and Inclusive AI, a non-binding international statement adopted on 21 May 2024 by 10 countries plus the European Union at…

AI Policy & Regulation

SimpleQA

SimpleQA is a factuality benchmark released by OpenAI on October 30, 2024 that measures whether large language models can answer short, fact-seeking questions correctly instead of producing hallucinations.

AI BenchmarksNatural Language Processing

Sleeper Agents (paper)

"Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training" is a January 2024 Anthropic research paper (arXiv:2401.05566, submitted 10 January 2024) that deliberately trained large language…

AI ResearchAnthropic

Stuart Russell

Stuart Jonathan Russell (born 1962) is a British computer scientist, professor of computer science at the University of California, Berkeley, and one of the most influential figures in modern artificial…

Computer SciencePeople

Studio Ghibli ChatGPT moment

The "Studio Ghibli ChatGPT moment" refers to a viral internet trend that erupted in late March 2025, after OpenAI released a new native image generation capability inside ChatGPT, powered by the GPT-4o model.

AI Companies

Superalignment

Superalignment is the technical problem of steering and controlling AI systems that are far more capable than their human supervisors, that is, systems at or beyond the level of superintelligence, and the name…

AI Alignment

Superintelligence

Superintelligence is a hypothetical form of artificial intelligence that surpasses all human cognitive abilities across virtually every domain, including scientific reasoning, social skills, creativity, and…

Artificial Intelligence

SynthID

SynthID is a family of digital watermarking technologies developed by Google DeepMind for marking and identifying content generated by generative AI systems.

Generative AIGoogle DeepMind

Tilly Norwood

Tilly Norwood is a fully AI-generated synthetic performer, marketed as an "AI actress," created in 2025 by the Dutch comedian, actress, and producer Eline Van der Velden through her production company…

AI Companies

Timnit Gebru

Timnit Gebru is an Ethiopian-born computer scientist and a leading researcher in AI ethics, best known for co-authoring the 2018 "Gender Shades" study on bias in facial recognition, co-leading Google's Ethical…

AI EthicsPeople

Toby Ord

Toby Ord is an Australian moral philosopher at the University of Oxford who founded the effective-altruism organisation Giving What We Can in 2009 and wrote the 2020 book The Precipice: Existential Risk and…

AI EthicsPeople

TopK SAE

A TopK SAE (TopK sparse autoencoder) is a variant of sparse autoencoder that enforces sparsity by keeping only the K largest latent pre-activations for each input and zeroing all the rest

Machine Learning

Transhumanism

Transhumanism is an intellectual and cultural movement that holds that the human condition can and should be fundamentally improved through science and technology, in particular through technologies that…

AI EthicsAI History

Tristan Harris

Tristan Harris is an American technology ethicist, public advocate, and entrepreneur best known as a co-founder and the president of the Center for Humane Technology

People

WMDP benchmark

The Weapons of Mass Destruction Proxy (WMDP) is a publicly released benchmark and unlearning testbed for large language models, introduced in the paper "The WMDP Benchmark: Measuring and Reducing Malicious Use…

AI Benchmarks

William MacAskill

William David MacAskill (born William Crouch; 24 March 1987) is a Scottish moral philosopher, author, and a co-founder of the effective altruism movement, best known as the leading public proponent of…

AI EthicsPeople

Youth AI Safety Institute

The Youth AI Safety Institute is an independent research and testing organization launched by Common Sense Media on May 5, 2026, in San Francisco to evaluate the artificial intelligence products that children…

Research Organizations