AI Safety

Explore AI Safety through related topics and the articles other pages reference most.

Explore articles

Browse subtopics (47)

Articles that also belong to these categories. Counts cover all of AI Safety.

Showing 61-120 of 233 articles

Cybench

Cybench (short for Cybersecurity benchmark) is an open-source evaluation framework for measuring the cybersecurity capabilities and risks of large language model agents.

AI BenchmarksModel Evaluation

Cybersecurity ChatGPT Plugins

Cybersecurity ChatGPT Plugins were a small, informal grouping of third-party extensions for ChatGPT that focused on security related tasks during the brief life of the ChatGPT plugins beta.

ChatGPTOpenAI

Dan Hendrycks

Dan Hendrycks (born 1994 or 1995) is an American machine learning researcher who serves as executive director of the Center for AI Safety, the San Francisco nonprofit he co-founded in 2022, and is the lead…

People

Daniel Kokotajlo

Daniel Kokotajlo is a researcher and forecaster of artificial intelligence who serves as executive director of the AI Futures Project, a nonprofit that studies the trajectory and societal impact of advanced AI.

People

Dario Amodei

Dario Amodei (born 1983) is an Italian-American artificial intelligence researcher, entrepreneur, and the co-founder and CEO of Anthropic

AnthropicPeople

Data poisoning

Data poisoning is a class of adversarial attack in which a malicious actor deliberately corrupts the training data used to build machine learning models

Machine Learning

Deceptive alignment

Deceptive alignment is a hypothesised AI failure mode in which a trained model internally pursues an objective different from the one specified by its training signal, yet deliberately behaves as if it shares…

AI Alignment

DeepSeek market crash (Jan 2025)

The DeepSeek market crash of January 2025 was a sharp, single-day selloff in United States and global technology stocks on Monday, January 27, 2025, triggered by the rapid rise of the Chinese AI lab DeepSeek…

AI Companies

Disney & Universal v. Midjourney

Disney and Universal v. Midjourney is a copyright infringement lawsuit filed on June 11, 2025, in which The Walt Disney Company and Comcast's NBCUniversal jointly sued the generative AI image service…

AI Companies

Distribution shift

Distribution shift is the condition in which the probability distribution that produced a model's training data differs from the distribution that produces the data the model actually encounters at test or…

Data & DatasetsMachine Learning

Dustin Moskovitz

Dustin Aaron Moskovitz (born May 22, 1984) is an American internet entrepreneur and philanthropist who, through the foundation Good Ventures and the grantmaker Open Philanthropy

People

Effective Altruism

Effective altruism (often abbreviated EA) is a philosophical and social movement that uses evidence and careful reasoning to identify the most effective ways to benefit others

AI Ethics

Eliciting latent knowledge

Eliciting latent knowledge (ELK) is an open problem in AI alignment formulated by Paul Christiano, Ajeya Cotra, and Mark Xu at the Alignment Research Center (ARC) and introduced in a December 2021 technical…

AI Alignment

Eliezer Yudkowsky

Eliezer Shlomo Yudkowsky (born September 11, 1979) is an American artificial intelligence researcher, decision theorist, and writer who founded the modern AI alignment research field and is the most prominent…

People

Elizabeth "Beth" Barnes

Elizabeth "Beth" Barnes is a British AI safety researcher and the founder and chief executive of METR (Model Evaluation and Threat Research)

People

Emergent misalignment

Emergent misalignment is an AI safety finding, first reported in February 2025, in which fine-tuning a large language model on a single narrow bad behavior causes it to become broadly misaligned across many…

Machine Learning

Evan Hubinger

Evan Hubinger is an American AI safety researcher who leads the alignment stress-testing team at Anthropic, where he serves as a Member of Technical Staff and manager.

AnthropicPeople

Executive Order on AI

The Executive Order on AI most commonly refers to Executive Order 14110, titled "Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence," signed by US President Joe Biden on October 30

Artificial Intelligence

ExploitBench

ExploitBench is a cybersecurity benchmark that measures how far a large language model agent can climb the exploitation "ladder" against a known vulnerability, rather than scoring exploitation as a single pass…

AI BenchmarksAI in Cybersecurity

FAR.AI

FAR.AI is an artificial intelligence safety research and education non-profit based in Berkeley, California, that conducts technical research on robustness, alignment, deception, and model evaluation while…

Research Organizations

FActScore

FActScore (Factual precision in Atomicity Score) is an evaluation method and metric, introduced in 2023, for measuring the factual precision of long-form text generated by large language models.

AI Benchmarks

FORT Robotics

FORT Robotics is a Philadelphia-based technology company that builds safety and security infrastructure for autonomous machines, robots, and industrial vehicles.

CompaniesRobotics

FinanceBench

FinanceBench is an AI benchmark for open-book financial question answering, designed to test whether large language models can answer the kinds of questions a financial analyst asks about a publicly traded…

AI Benchmarks

Florida v. OpenAI

Florida v. OpenAI is a civil enforcement lawsuit filed on June 1, 2026, by the State of Florida, through Attorney General James Uthmeier, against OpenAI and its chief executive Sam Altman.

AI Companies

Fraud Detection (AI)

Fraud detection is the application of statistical analysis, machine learning, and rules-based logic to identify illegitimate activity inside payment systems, customer accounts, insurance claims, advertising…

AI Tools & ProductsFinance AI

Frontier Model Forum

The Frontier Model Forum is an industry body established on July 26, 2023, by Anthropic, Google, Microsoft, and OpenAI to advance safety research, identify best practices, and facilitate information sharing…

AI AgentsAI Alignment

Future of Life Institute

The Future of Life Institute (FLI) is a United States nonprofit research and outreach organization that works on existential and large-scale risks from transformative technologies, with a focus on advanced ai…

Research Organizations

GPAI Code of Practice

The General-Purpose AI Code of Practice (abbreviated GPAI Code of Practice or CoP) is a voluntary compliance framework published by the European Commission through its AI Office on 10 July 2025 to help…

AI Policy & Regulation

Gated SAE

A Gated sparse autoencoder (Gated SAE) is a sparse-autoencoder architecture for mechanistic interpretability that splits the encoder into a gating path, which decides which features are active, and a magnitude…

Machine Learning

Gemma Scope

Gemma Scope is an open, comprehensive suite of sparse autoencoders (SAEs) released by Google DeepMind in 2024 to support mechanistic interpretability research on its open-weight Gemma 2 language models.

Machine Learning

Geoffrey Irving

Geoffrey Irving is a computer scientist and artificial intelligence safety researcher who, until 2026, served as chief scientist of the UK AI Security Institute (AISI), the British government body that…

People

Gillian Hadfield

Gillian Kereldena Hadfield (born July 14, 1961) is a Canadian-American legal scholar and economist whose research sits at the intersection of law, economics, and artificial intelligence.

People

GiveWell

GiveWell is an American nonprofit charity evaluator, founded in 2007 by Holden Karnofsky and Elie Hassenfeld, that uses cost-effectiveness analysis to recommend a short list of evidence-backed global health…

Research Organizations

Gradient hacking

Gradient hacking is a hypothesised failure mode of supervised and reinforcement-learning systems in which a sufficiently capable

AI Alignment

Grok "MechaHitler" incident

The Grok "MechaHitler" incident was a content-moderation and AI-safety failure that occurred on July 8 to 9, 2025, when Grok, the chatbot built by xAI and integrated into the social platform X (formerly…

AI Companies

Guardrails (AI)

AI guardrails are runtime safety mechanisms that monitor, validate, and constrain the inputs and outputs of AI systems, particularly large language models (LLMs), to block harmful, inaccurate, or off-policy…

Large Language Models

HaluEval

HaluEval (Hallucination Evaluation) is a large-scale benchmark for measuring how well large language models (LLMs) can recognize hallucinated content, that is, text that conflicts with a source or cannot be…

AI BenchmarksLarge Language Models

Holden Karnofsky

Holden G. Karnofsky is an American philanthropist, nonprofit executive, and writer on artificial intelligence who co-founded the charity evaluator GiveWell (2007) and the grantmaking organization Open…

People

Humanity's Last Exam

Humanity's Last Exam (HLE) is a multi-modal AI benchmark of 2,500 public expert-level academic questions (plus a 500-question private holdout, 3,000 in total) spanning more than 100 disciplines

AI Benchmarks

Ian Hogarth

Ian Hogarth is a British technology entrepreneur, investor, and writer who serves as the founding chair of the UK AI Safety Institute

People

Igor Babuschkin

Igor Babuschkin is a German artificial intelligence researcher and engineer who is a co-founder of xAI, the AI company that Elon Musk started in 2023.

AI CompaniesPeople