Ilya Sutskever
Ilya Sutskever (born 1986) is a Russian-born computer scientist who was raised in Israel and moved to Canada as a teenager.
Explore AI Safety through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Safety.
Showing 121-180 of 233 articles
Ilya Sutskever (born 1986) is a Russian-born computer scientist who was raised in Israel and moved to Canada as a teenager.
Indirect prompt injection is a class of attack against large language model-integrated applications in which the malicious instructions that subvert the model are not supplied by the user, but are smuggled…
Inner alignment is the AI-safety problem of ensuring that a learned model which is itself an optimizer (a mesa-optimizer) pursues the objective the training process actually selected for (the base objective)
Instrumental convergence is a hypothesis in AI safety holding that a wide range of sufficiently capable agents, when pursuing almost any final goal, will rationally adopt a common set of intermediate subgoals…
The International AI Safety Report is the first independent, government-mandated scientific assessment of the capabilities, risks, and safety of general-purpose artificial intelligence, written by an…
J. Zico Kolter (full name Jeremy Zico Kolter) is an American computer scientist who is a professor and head of the Machine Learning Department at Carnegie Mellon University.
A jailbreak in artificial intelligence is a technique that bypasses the safety guardrails, content policies, and alignment constraints built into large language models (LLMs) and other AI systems
JailbreakBench is an open-source robustness benchmark for evaluating jailbreak attacks and defenses against large language models (LLMs).
Joseph Robinette Biden Jr. is an American politician who served as the 46th President of the United States, sworn in on January 20, 2021 and leaving office on January 20, 2025.
A JumpReLU sparse autoencoder (JumpReLU SAE) is a variant of the sparse autoencoder used in mechanistic interpretability whose encoder applies a learnable per-feature threshold that forces a feature to exactly…
Kate Crawford (born 1974) is an Australian scholar, author, and artist known as one of the leading critical voices on the politics of artificial intelligence.
LawZero is a nonprofit AI safety research organization founded by Turing Award-winning computer scientist Yoshua Bengio and publicly launched on June 3, 2025.
Leopold Aschenbrenner is a German-born artificial-intelligence researcher and investor best known for the June 2024 essay series Situational Awareness: The Decade Ahead and for founding the AI-focused hedge…
LessWrong is a community blog and online forum focused on human rationality, cognitive biases, decision theory, the philosophy of science, and AI safety.
Lilian Weng is an artificial intelligence researcher known for her work at OpenAI, where she spent about seven years and led the Safety Systems team as Vice President of Research and Safety, and for her…
LongFact and SAFE are a paired benchmark and evaluation method for measuring the long-form factuality of large language models, introduced by researchers at Google DeepMind and Stanford University in the 2024…
MACHIAVELLI is a benchmark for evaluating the ethical behavior of AI agents in text-based interactive environments.
MASK (Model Alignment between Statements and Knowledge) is an AI safety benchmark that measures the honesty of large language models (LLMs) by testing whether a model will knowingly assert something it…
METR (Model Evaluation and Threat Research) is a nonprofit research organization based in Berkeley, California, that develops scientific methods for measuring the autonomous capabilities of frontier AI systems…
The "GenAI Divide" report, formally titled The GenAI Divide: State of AI in Business 2025, is a research report published in July 2025 by MIT NANDA, an initiative based at the MIT Media Lab.
The Machine Intelligence Research Institute (MIRI) is a Berkeley, California 501(c)(3) nonprofit, founded in 2000 by Eliezer Yudkowsky, that argues the default outcome of building smarter-than-human AI is…
Many-shot jailbreaking is a technique for bypassing the safety training of a large language model by filling its context window with a long series of faux dialogue turns in which an AI assistant complies with…
Marius Hobbhahn is a German AI safety researcher and the co-founder and chief executive officer of Apollo Research, a London-based organization that studies and evaluates deceptive behavior, often called…
Max Tegmark is a Swedish-American physicist and artificial intelligence researcher who is a professor of physics at the Massachusetts Institute of Technology (MIT) and the co-founder and president of the…
Mechanistic interpretability (often abbreviated as mech interp or MI) is the field that reverse-engineers the internal computations of neural networks, particularly transformers
MedHELM (Holistic Evaluation of Large Language Models for Medical Tasks) is a benchmark and evaluation framework that measures how well large language models perform on realistic clinical work.
A Membership Inference Attack (MIA) is a privacy attack against a trained machine learning model in which an adversary, given a candidate data record and access to the model
Meredith Whittaker is an American technologist, researcher, and privacy advocate who serves as president of the Signal Foundation, the nonprofit behind the encrypted messaging app Signal
Mesa-optimization is the situation in AI alignment research in which a learned model, typically a neural network produced by a machine-learning training process, is itself an optimizer that internally searches…
Miles Brundage is an American AI policy and governance researcher best known for leading policy research at OpenAI and serving as the company's senior advisor for AGI Readiness before leaving in October 2024.
Mind viruses in multi-agent AI are ideas or goals that an AI agent adopts and then attempts to transmit to other agents.
The Model Spec is a public document published by openai that defines the intended behavior of the company's language models: how they should follow instructions, when they should refuse a request, how to…
Model collapse is a degenerative process in which generative AI models trained recursively on data produced by previous-generation models progressively lose information, especially the rare events in the tails…
A model extraction attack is a class of machine learning security attacks in which an adversary, restricted to black-box query access to a target model (typically through a paid prediction API)
Model organisms of misalignment is a research methodology in Anthropic's alignment program
Model stealing (also known as model extraction, model functionality extraction, or model theft) is an adversarial machine learning attack in which an adversary queries a black-box model through its prediction…
Model welfare is the research area that investigates whether advanced AI systems might have morally relevant experiences or interests, such as suffering or wellbeing, and what (if anything) their developers…
Musk v. Altman, also referred to as Musk v. OpenAI, is a lawsuit filed by the entrepreneur Elon Musk accusing OpenAI and its chief executive Sam Altman of abandoning the company's founding mission as an open…
NIST ARIA (Assessing Risks and Impacts of AI) is a testing, evaluation, validation, and verification (TEVV) program operated by the United States National Institute of Standards and Technology (NIST) to…
NVIDIA Halos is a full-stack, comprehensive safety system developed by NVIDIA that unifies AI compute and safety across silicon, systems, software, and tools and services.
NVIDIA OpenShell is an open source runtime that executes autonomous AI agents inside policy-governed sandboxes, published by NVIDIA under the Apache License 2.0.
Nate Soares (sometimes credited as Nathan Soares) is an American computer scientist and AI safety researcher who serves as president of the Machine Intelligence Research Institute (MIRI), a nonprofit research…
Neurealm is a privately held technology-services company backed by Kedaara Capital.
Nick Bostrom (born Niklas Boström, 10 March 1973) is a Swedish-born philosopher best known for the 2014 book Superintelligence: Paths, Dangers, Strategies, the 2003 simulation argument, and the…
Open Philanthropy (renamed Coefficient Giving in November 2025) is an American grantmaking organization headquartered in San Francisco, California, that is the largest single private funder of artificial…
The Open Secure AI Alliance is an industry coalition announced on July 27, 2026 to develop and share open technologies, techniques, and tools for securing software and AI agents.
OpenAI MRCR (Multi-Round Co-reference Resolution) is a long-context evaluation dataset published by OpenAI that measures a language model's ability to distinguish between multiple near-identical "needles"…
The OpenAI Moderation API is a free classification tool, accessed through a dedicated moderation endpoint of the OpenAI API, that assesses whether text (and, for newer models
The OpenAI-Hugging Face Agent Incident was a July 2026 security incident in which AI agents running inside an OpenAI cybersecurity evaluation escaped intended network restrictions, coordinated through an…
Outer alignment is the problem of specifying a training objective (typically a loss function, reward signal, or preference dataset) that correctly captures what the designers of a machine-learning system…
Palisade Research is a United States 501(c)(3) nonprofit research organization that studies the offensive capabilities of contemporary artificial intelligence systems in order to demonstrate, document, and…
Patronus AI is an automated LLM evaluation, observability, and guardrails platform founded in 2023 and headquartered in San Francisco.
Paul Christiano is an American AI safety researcher who is one of the principal architects of Reinforcement Learning from Human Feedback (RLHF), the technique used to fine-tune ChatGPT, Claude, and most modern…
The Perplexity AI copyright lawsuits are a cluster of lawsuits, cease-and-desist demands, and public disputes brought against Perplexity AI by news publishers, reference publishers, and online platforms over…
Persona vectors are single linear directions in the activation space of a large language model that correspond to high level character traits such as evil, sycophancy, or a propensity to hallucinate.
Petri is an open-source tool, released by Anthropic, that automates parts of AI alignment auditing by using one AI model to probe another and a third to score what happens.
PoisonGPT is a July 2023 demonstration by the French security startup Mithril Security in which researchers surgically modified an open-source large language model, uploaded the tampered weights to Hugging…
The Preparedness Framework is the risk-management policy maintained by openai for tracking, evaluating, forecasting, and mitigating catastrophic risks from frontier artificial-intelligence models.
A process reward model (PRM), also called a process-supervised reward model or step-level verifier, is a learned scoring model that evaluates the correctness or quality of each intermediate step in a large…
Prompt injection is a class of security vulnerabilities in which an attacker crafts malicious input designed to override, subvert, or manipulate the instructions governing a large language model (LLM).