AI Safety

Explore AI Safety through related topics and the articles other pages reference most.

Explore articles

Browse subtopics (47)

Articles that also belong to these categories. Counts cover all of AI Safety.

Showing 121-180 of 233 articles

Indirect prompt injection

Indirect prompt injection is a class of attack against large language model-integrated applications in which the malicious instructions that subvert the model are not supplied by the user, but are smuggled…

Large Language Models

Inner alignment

Inner alignment is the AI-safety problem of ensuring that a learned model which is itself an optimizer (a mesa-optimizer) pursues the objective the training process actually selected for (the base objective)

AI Alignment

Instrumental convergence

Instrumental convergence is a hypothesis in AI safety holding that a wide range of sufficiently capable agents, when pursuing almost any final goal, will rationally adopt a common set of intermediate subgoals…

AI Alignment

J. Zico Kolter

J. Zico Kolter (full name Jeremy Zico Kolter) is an American computer scientist who is a professor and head of the Machine Learning Department at Carnegie Mellon University.

People

Joe Biden

Joseph Robinette Biden Jr. is an American politician who served as the 46th President of the United States, sworn in on January 20, 2021 and leaving office on January 20, 2025.

AI Policy & RegulationPeople

JumpReLU SAE

A JumpReLU sparse autoencoder (JumpReLU SAE) is a variant of the sparse autoencoder used in mechanistic interpretability whose encoder applies a learnable per-feature threshold that forces a feature to exactly…

Machine Learning

Kate Crawford

Kate Crawford (born 1974) is an Australian scholar, author, and artist known as one of the leading critical voices on the politics of artificial intelligence.

People

LawZero

LawZero is a nonprofit AI safety research organization founded by Turing Award-winning computer scientist Yoshua Bengio and publicly launched on June 3, 2025.

Research Organizations

Leopold Aschenbrenner

Leopold Aschenbrenner is a German-born artificial-intelligence researcher and investor best known for the June 2024 essay series Situational Awareness: The Decade Ahead and for founding the AI-focused hedge…

People

LessWrong

LessWrong is a community blog and online forum focused on human rationality, cognitive biases, decision theory, the philosophy of science, and AI safety.

Lilian Weng

Lilian Weng is an artificial intelligence researcher known for her work at OpenAI, where she spent about seven years and led the Safety Systems team as Vice President of Research and Safety, and for her…

Machine LearningPeople

LongFact / SAFE

LongFact and SAFE are a paired benchmark and evaluation method for measuring the long-form factuality of large language models, introduced by researchers at Google DeepMind and Stanford University in the 2024…

AI Benchmarks

MASK

MASK (Model Alignment between Statements and Knowledge) is an AI safety benchmark that measures the honesty of large language models (LLMs) by testing whether a model will knowingly assert something it…

AI Benchmarks

METR

METR (Model Evaluation and Threat Research) is a nonprofit research organization based in Berkeley, California, that develops scientific methods for measuring the autonomous capabilities of frontier AI systems…

AI BenchmarksResearch Organizations

MIT "GenAI Divide" report (2025)

The "GenAI Divide" report, formally titled The GenAI Divide: State of AI in Business 2025, is a research report published in July 2025 by MIT NANDA, an initiative based at the MIT Media Lab.

AI Companies

Many-shot jailbreaking

Many-shot jailbreaking is a technique for bypassing the safety training of a large language model by filling its context window with a long series of faux dialogue turns in which an AI assistant complies with…

AnthropicLarge Language Models

Marius Hobbhahn

Marius Hobbhahn is a German AI safety researcher and the co-founder and chief executive officer of Apollo Research, a London-based organization that studies and evaluates deceptive behavior, often called…

People

Max Tegmark

Max Tegmark is a Swedish-American physicist and artificial intelligence researcher who is a professor of physics at the Massachusetts Institute of Technology (MIT) and the co-founder and president of the…

AI EthicsPeople

Mechanistic interpretability

Mechanistic interpretability (often abbreviated as mech interp or MI) is the field that reverse-engineers the internal computations of neural networks, particularly transformers

Interpretability

MedHELM

MedHELM (Holistic Evaluation of Large Language Models for Medical Tasks) is a benchmark and evaluation framework that measures how well large language models perform on realistic clinical work.

AI Benchmarks

Membership Inference Attack

A Membership Inference Attack (MIA) is a privacy attack against a trained machine learning model in which an adversary, given a candidate data record and access to the model

Machine Learning

Meredith Whittaker

Meredith Whittaker is an American technologist, researcher, and privacy advocate who serves as president of the Signal Foundation, the nonprofit behind the encrypted messaging app Signal

AI EthicsPeople

Mesa-optimization

Mesa-optimization is the situation in AI alignment research in which a learned model, typically a neural network produced by a machine-learning training process, is itself an optimizer that internally searches…

AI Alignment

Miles Brundage

Miles Brundage is an American AI policy and governance researcher best known for leading policy research at OpenAI and serving as the company's senior advisor for AGI Readiness before leaving in October 2024.

AI Policy & RegulationPeople

Model Spec

The Model Spec is a public document published by openai that defines the intended behavior of the company's language models: how they should follow instructions, when they should refuse a request, how to…

AI AlignmentOpenAI

Model collapse

Model collapse is a degenerative process in which generative AI models trained recursively on data produced by previous-generation models progressively lose information, especially the rare events in the tails…

Machine Learning

Model extraction attack

A model extraction attack is a class of machine learning security attacks in which an adversary, restricted to black-box query access to a target model (typically through a paid prediction API)

Machine Learning

Model stealing

Model stealing (also known as model extraction, model functionality extraction, or model theft) is an adversarial machine learning attack in which an adversary queries a black-box model through its prediction…

Machine Learning

Model welfare

Model welfare is the research area that investigates whether advanced AI systems might have morally relevant experiences or interests, such as suffering or wellbeing, and what (if anything) their developers…

AI Ethics

Musk v. Altman (Musk v. OpenAI)

Musk v. Altman, also referred to as Musk v. OpenAI, is a lawsuit filed by the entrepreneur Elon Musk accusing OpenAI and its chief executive Sam Altman of abandoning the company's founding mission as an open…

AI Companies

NIST ARIA

NIST ARIA (Assessing Risks and Impacts of AI) is a testing, evaluation, validation, and verification (TEVV) program operated by the United States National Institute of Standards and Technology (NIST) to…

AI Policy & RegulationModel Evaluation

NVIDIA Halos

NVIDIA Halos is a full-stack, comprehensive safety system developed by NVIDIA that unifies AI compute and safety across silicon, systems, software, and tools and services.

HardwareNVIDIA

NVIDIA OpenShell

NVIDIA OpenShell is an open source runtime that executes autonomous AI agents inside policy-governed sandboxes, published by NVIDIA under the Apache License 2.0.

AI AgentsDeveloper Tools

Nate Soares

Nate Soares (sometimes credited as Nathan Soares) is an American computer scientist and AI safety researcher who serves as president of the Machine Intelligence Research Institute (MIRI), a nonprofit research…

People

Nick Bostrom

Nick Bostrom (born Niklas Boström, 10 March 1973) is a Swedish-born philosopher best known for the 2014 book Superintelligence: Paths, Dangers, Strategies, the 2003 simulation argument, and the…

AI EthicsPeople

Open Philanthropy

Open Philanthropy (renamed Coefficient Giving in November 2025) is an American grantmaking organization headquartered in San Francisco, California, that is the largest single private funder of artificial…

PeopleResearch Organizations

OpenAI Moderation API

The OpenAI Moderation API is a free classification tool, accessed through a dedicated moderation endpoint of the OpenAI API, that assesses whether text (and, for newer models

Developer ToolsOpenAI

Outer alignment

Outer alignment is the problem of specifying a training objective (typically a loss function, reward signal, or preference dataset) that correctly captures what the designers of a machine-learning system…

AI Alignment

Palisade Research

Palisade Research is a United States 501(c)(3) nonprofit research organization that studies the offensive capabilities of contemporary artificial intelligence systems in order to demonstrate, document, and…

Research Organizations

Paul Christiano

Paul Christiano is an American AI safety researcher who is one of the principal architects of Reinforcement Learning from Human Feedback (RLHF), the technique used to fine-tune ChatGPT, Claude, and most modern…

People

Perplexity AI copyright lawsuits

The Perplexity AI copyright lawsuits are a cluster of lawsuits, cease-and-desist demands, and public disputes brought against Perplexity AI by news publishers, reference publishers, and online platforms over…

AI Companies

PoisonGPT

PoisonGPT is a July 2023 demonstration by the French security startup Mithril Security in which researchers surgically modified an open-source large language model, uploaded the tampered weights to Hugging…

Preparedness Framework (OpenAI)

The Preparedness Framework is the risk-management policy maintained by openai for tracking, evaluating, forecasting, and mitigating catastrophic risks from frontier artificial-intelligence models.

OpenAI

Prompt injection

Prompt injection is a class of security vulnerabilities in which an attacker crafts malicious input designed to override, subvert, or manipulate the instructions governing a large language model (LLM).

Large Language Models