Purple Llama
Purple Llama is an umbrella project from Meta that bundles open, permissively licensed trust and safety tools and evaluations for generative AI.
Explore AI Safety through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Safety.
Showing 181-233 of 233 articles
Purple Llama is an umbrella project from Meta that bundles open, permissively licensed trust and safety tools and evaluations for generative AI.
Reinforcement Learning from AI Feedback (RLAIF) is a family of alignment techniques for large language models in which the preference labels used to fine-tune a model are produced by another AI system
Raine v. OpenAI, Inc. is a wrongful-death lawsuit filed on August 26, 2025, in the Superior Court of California for the County of San Francisco (case number CGC-25-628528) by Matthew and Maria Raine, the…
Recursive reward modeling (RRM) is a proposed approach to the scalable oversight problem in AI alignment
Recursive self-improvement (RSI) is a process in which an artificial intelligence system improves its own intelligence or its ability to improve itself, so that each enhancement increases its capacity for…
Red teaming in artificial intelligence is the systematic, adversarial testing of an AI system to find vulnerabilities, biases, harmful outputs, and other failure modes before deployment or as part of ongoing…
Redwood Research is a nonprofit AI safety organization founded in 2021 and headquartered in Berkeley, California, best known for pioneering the "AI control" research paradigm and for its landmark December 2024…
The Reflection 70B controversy was an open-source AI credibility episode that began on September 5, 2024, when Matt Shumer, co-founder and chief executive of OthersideAI (the company behind the AI writing…
The refusal direction is a finding from mechanistic interpretability research that the refusal behavior of safety fine-tuned chat language models is mediated by a single
The Remote Labor Index (RLI) is an AI benchmark that measures how well AI agents can complete real, paid remote knowledge work end to end.
Representation Engineering (often abbreviated RepE) is a top-down approach to artificial-intelligence transparency and control that reads and manipulates high-level concepts (such as honesty, harmlessness, and…
Responsible AI (RAI) is a framework for developing, deploying, and governing artificial intelligence systems in ways that are ethical, transparent, accountable, and aligned with human values.
A Responsible Scaling Policy (RSP) is a self-imposed governance framework in which a frontier AI developer commits in advance to safety practices, capability evaluations, deployment restrictions, and security…
Reward hacking (also called specification gaming) is a failure mode in artificial intelligence in which a system maximizes its given objective or reward signal through unintended shortcuts, exploits, or…
Robot safety is the discipline concerned with minimizing the risk of physical harm, property damage, and other hazards arising from the operation of robotic systems.
Roman Vladimirovich Yampolskiy (born August 13, 1979) is a Latvian-American computer scientist and a tenured associate professor in the Department of Computer Engineering and Computer Science at the University…
Rule-Based Rewards (RBR) is a safety-alignment technique introduced by OpenAI in July 2024 that replaces large quantities of human-labeled safety preference data with an explicit collection of natural-language…
Rumman Chowdhury (born April 1, 1980) is an American data scientist and social scientist known for her work in responsible AI, algorithmic auditing, and the study of algorithmic bias.
SB 1047, officially the Safe and Secure Innovation for Frontier Artificial Intelligence Models Act
Sabotage evaluations are a set of tests, introduced by Anthropic in October 2024, that try to measure whether a frontier language model is capable of covertly subverting human oversight, decision-making, and…
Safe Superintelligence Inc. (abbreviated SSI) is an AI safety research company that is building a single product, a safe superintelligence, and nothing else.
Sam Bankman-Fried (born 1992), often called SBF, is an American former cryptocurrency executive who founded the trading firm Alameda Research and the FTX exchange.
Sandbagging, in the context of AI safety, refers to the strategic and intentional underperformance of an AI system on a capability evaluation or specific task, typically to hide a capability from human…
Scalable oversight is the AI safety problem of how humans can reliably supervise, evaluate, and provide training signal to artificial intelligence systems whose capabilities approach, equal, or exceed those of…
The Seoul Declaration is the short name for the Seoul Declaration for Safe, Innovative and Inclusive AI, a non-binding international statement adopted on 21 May 2024 by 10 countries plus the European Union at…
ShieldGemma is a family of open safety-classifier models from Google, built on the Gemma family of lightweight open models.
SimpleQA is a factuality benchmark released by OpenAI on October 30, 2024 that measures whether large language models can answer short, fact-seeking questions correctly instead of producing hallucinations.
Situational Awareness: The Decade Ahead is a 165-page essay series published on June 4, 2024 by Leopold Aschenbrenner, a former member of OpenAI's Superalignment team.
"Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training" is a January 2024 Anthropic research paper (arXiv:2401.05566, submitted 10 January 2024) that deliberately trained large language…
Specification gaming is the phenomenon in which an optimizer satisfies the literal specification of an objective without producing the outcome that the designer actually wanted.
StrongREJECT is a benchmark and automated evaluator for measuring how well jailbreaking attacks actually work against large language models.
Stuart Jonathan Russell (born 1962) is a British computer scientist, professor of computer science at the University of California, Berkeley, and one of the most influential figures in modern artificial…
The "Studio Ghibli ChatGPT moment" refers to a viral internet trend that erupted in late March 2025, after OpenAI released a new native image generation capability inside ChatGPT, powered by the GPT-4o model.
Superalignment is the technical problem of steering and controlling AI systems that are far more capable than their human supervisors, that is, systems at or beyond the level of superintelligence, and the name…
Superintelligence is a hypothetical form of artificial intelligence that surpasses all human cognitive abilities across virtually every domain, including scientific reasoning, social skills, creativity, and…
Sycophancy in artificial intelligence is the tendency of large language models to tell users what they want to hear: tailoring responses to match a user's perceived beliefs, preferences, or emotional state…
SynthID is a family of digital watermarking technologies developed by Google DeepMind for marking and identifying content generated by generative AI systems.
A system prompt is a special set of instructions, guidelines, persona definitions, and contextual information given to a large language model (LLM) before any user input
The task-completion time horizon is a metric for AI capability proposed by METR that expresses a model's ability in units of human time: it is the length of task
The Anthropic Institute is a research organization within Anthropic dedicated to studying the societal challenges posed by increasingly powerful artificial intelligence systems.
Tilly Norwood is a fully AI-generated synthetic performer, marketed as an "AI actress," created in 2025 by the Dutch comedian, actress, and producer Eline Van der Velden through her production company…
Timnit Gebru is an Ethiopian-born computer scientist and a leading researcher in AI ethics, best known for co-authoring the 2018 "Gender Shades" study on bias in facial recognition, co-leading Google's Ethical…
Toby Ord is an Australian moral philosopher at the University of Oxford who founded the effective-altruism organisation Giving What We Can in 2009 and wrote the 2020 book The Precipice: Existential Risk and…
A TopK SAE (TopK sparse autoencoder) is a variant of sparse autoencoder that enforces sparsity by keeping only the K largest latent pre-activations for each input and zeroing all the rest
ToxiGen is a large-scale, machine-generated dataset designed for adversarial and implicit hate speech detection.
Transhumanism is an intellectual and cultural movement that holds that the human condition can and should be fundamentally improved through science and technology, in particular through technologies that…
Tristan Harris is an American technology ethicist, public advocate, and entrepreneur best known as a co-founder and the president of the Center for Humane Technology
TruthfulQA is a benchmark designed to measure whether large language models (LLMs) generate truthful answers to questions.
The UK AI Security Institute (AISI) is a research organization within the United Kingdom's Department for Science, Innovation and Technology (DSIT) that conducts pre-deployment evaluations of frontier AI…
The US AI Safety Institute (USAISI or US AISI), since June 2025 the Center for AI Standards and Innovation (CAISI), is a research and evaluation body housed within the National Institute of Standards and…
The Weapons of Mass Destruction Proxy (WMDP) is a publicly released benchmark and unlearning testbed for large language models, introduced in the paper "The WMDP Benchmark: Measuring and Reducing Malicious Use…
William David MacAskill (born William Crouch; 24 March 1987) is a Scottish moral philosopher, author, and a co-founder of the effective altruism movement, best known as the leading public proponent of…
The Youth AI Safety Institute is an independent research and testing organization launched by Common Sense Media on May 5, 2026, in San Francisco to evaluate the artificial intelligence products that children…