Cybench
Cybench (short for Cybersecurity benchmark) is an open-source evaluation framework for measuring the cybersecurity capabilities and risks of large language model agents.
Explore AI Safety through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Safety.
Showing 61-120 of 233 articles
Cybench (short for Cybersecurity benchmark) is an open-source evaluation framework for measuring the cybersecurity capabilities and risks of large language model agents.
Cybersecurity ChatGPT Plugins were a small, informal grouping of third-party extensions for ChatGPT that focused on security related tasks during the brief life of the ChatGPT plugins beta.
Dan Hendrycks (born 1994 or 1995) is an American machine learning researcher who serves as executive director of the Center for AI Safety, the San Francisco nonprofit he co-founded in 2022, and is the lead…
Daniel Kokotajlo is a researcher and forecaster of artificial intelligence who serves as executive director of the AI Futures Project, a nonprofit that studies the trajectory and societal impact of advanced AI.
Dario Amodei (born 1983) is an Italian-American artificial intelligence researcher, entrepreneur, and the co-founder and CEO of Anthropic
Data poisoning is a class of adversarial attack in which a malicious actor deliberately corrupts the training data used to build machine learning models
Deceptive alignment is a hypothesised AI failure mode in which a trained model internally pursues an objective different from the one specified by its training signal, yet deliberately behaves as if it shares…
The DeepSeek market crash of January 2025 was a sharp, single-day selloff in United States and global technology stocks on Monday, January 27, 2025, triggered by the rapid rise of the Chinese AI lab DeepSeek…
Dictionary learning, in the context of mechanistic interpretability, is the framework of decomposing the dense internal activations of a neural network into a sparse, weighted combination drawn from a large
Disney and Universal v. Midjourney is a copyright infringement lawsuit filed on June 11, 2025, in which The Walt Disney Company and Comcast's NBCUniversal jointly sued the generative AI image service…
Distribution shift is the condition in which the probability distribution that produced a model's training data differs from the distribution that produces the data the model actually encounters at test or…
Dustin Aaron Moskovitz (born May 22, 1984) is an American internet entrepreneur and philanthropist who, through the foundation Good Ventures and the grantmaker Open Philanthropy
The EU Action Plan on Cybersecurity and Artificial Intelligence is a European Commission Communication, published as COM(2026) 577 final and adopted in Strasbourg on 7 July 2026
Effective altruism (often abbreviated EA) is a philosophical and social movement that uses evidence and careful reasoning to identify the most effective ways to benefit others
Eliciting latent knowledge (ELK) is an open problem in AI alignment formulated by Paul Christiano, Ajeya Cotra, and Mark Xu at the Alignment Research Center (ARC) and introduced in a December 2021 technical…
Eliezer Shlomo Yudkowsky (born September 11, 1979) is an American artificial intelligence researcher, decision theorist, and writer who founded the modern AI alignment research field and is the most prominent…
Elizabeth "Beth" Barnes is a British AI safety researcher and the founder and chief executive of METR (Model Evaluation and Threat Research)
Emergent abilities are capabilities of large language models (LLMs) that are absent in smaller models but appear once a model reaches sufficient scale.
Emergent misalignment is an AI safety finding, first reported in February 2025, in which fine-tuning a large language model on a single narrow bad behavior causes it to become broadly misaligned across many…
Evan Hubinger is an American AI safety researcher who leads the alignment stress-testing team at Anthropic, where he serves as a Member of Technical Staff and manager.
Executive Order 14110, formally titled "Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence," was signed by President Joe Biden on October 30, 2023 and published in the Federal…
The Executive Order on AI most commonly refers to Executive Order 14110, titled "Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence," signed by US President Joe Biden on October 30
Existential risk from artificial intelligence (also called AI x-risk) is the hypothesis that the development of sufficiently advanced artificial intelligence could cause human extinction, permanent…
ExploitBench is a cybersecurity benchmark that measures how far a large language model agent can climb the exploitation "ladder" against a known vulnerability, rather than scoring exploitation as a single pass…
FAR.AI is an artificial intelligence safety research and education non-profit based in Berkeley, California, that conducts technical research on robustness, alignment, deception, and model evaluation while…
FActScore (Factual precision in Atomicity Score) is an evaluation method and metric, introduced in 2023, for measuring the factual precision of long-form text generated by large language models.
FORT Robotics is a Philadelphia-based technology company that builds safety and security infrastructure for autonomous machines, robots, and industrial vehicles.
FinanceBench is an AI benchmark for open-book financial question answering, designed to test whether large language models can answer the kinds of questions a financial analyst asks about a publicly traded…
Florida v. OpenAI is a civil enforcement lawsuit filed on June 1, 2026, by the State of Florida, through Attorney General James Uthmeier, against OpenAI and its chief executive Sam Altman.
Fraud detection is the application of statistical analysis, machine learning, and rules-based logic to identify illegitimate activity inside payment systems, customer accounts, insurance claims, advertising…
The Frontier Model Forum is an industry body established on July 26, 2023, by Anthropic, Google, Microsoft, and OpenAI to advance safety research, identify best practices, and facilitate information sharing…
Frontier models are artificial intelligence models at or near a selected boundary of capability, scale, or risk.
The Frontier Safety Framework (FSF) is Google DeepMind's risk-management framework for identifying and mitigating severe risks from advanced frontier AI models, first published on 17 May 2024 and updated to…
The Frontier Security Institute (FSI) is a Washington, D.C. organization launched by the Center for AI Safety (CAIS) to connect frontier artificial intelligence developers with the United States national…
The Future of Humanity Institute (FHI) was a multidisciplinary research institute at the University of Oxford that studied existential risk, the long-term future of humanity, and the transformative potential…
The Future of Life Institute (FLI) is a United States nonprofit research and outreach organization that works on existential and large-scale risks from transformative technologies, with a focus on advanced ai…
The General-Purpose AI Code of Practice (abbreviated GPAI Code of Practice or CoP) is a voluntary compliance framework published by the European Commission through its AI Office on 10 July 2025 to help…
GTG-1002 is the internal tracking name that Anthropic gave to a cyber espionage operation it says it detected in September 2025 and disclosed in November 2025
Garcia v. Character Technologies is a wrongful death and product liability lawsuit filed in late October 2024 in the United States District Court for the Middle District of Florida (Orlando Division)
A Gated sparse autoencoder (Gated SAE) is a sparse-autoencoder architecture for mechanistic interpretability that splits the encoder into a gating path, which decides which features are active, and a magnitude…
Gemma Scope is an open, comprehensive suite of sparse autoencoders (SAEs) released by Google DeepMind in 2024 to support mechanistic interpretability research on its open-weight Gemma 2 language models.
Geoffrey Irving is a computer scientist and artificial intelligence safety researcher who, until 2026, served as chief scientist of the UK AI Security Institute (AISI), the British government body that…
Gillian Kereldena Hadfield (born July 14, 1961) is a Canadian-American legal scholar and economist whose research sits at the intersection of law, economics, and artificial intelligence.
GiveWell is an American nonprofit charity evaluator, founded in 2007 by Holden Karnofsky and Elie Hassenfeld, that uses cost-effectiveness analysis to recommend a short list of evidence-backed global health…
Goodfire AI is a San Francisco-based artificial intelligence research lab and public benefit corporation focused on mechanistic interpretability
Gradient hacking is a hypothesised failure mode of supervised and reinforcement-learning systems in which a sufficiently capable
The Grok "MechaHitler" incident was a content-moderation and AI-safety failure that occurred on July 8 to 9, 2025, when Grok, the chatbot built by xAI and integrated into the social platform X (formerly…
Grok 3 jailbreak is the umbrella term for a class of reported guardrail-bypass findings against Grok 3, the third-generation large language model that xAI released on February 17, 2025.
The Grok child safety controversy was a multinational regulatory and legal crisis that began in late December 2025, when xAI's Grok chatbot and its image and video generator, Grok Imagine
Grounding in artificial intelligence is the process of anchoring an AI system's outputs to verifiable
AI guardrails are runtime safety mechanisms that monitor, validate, and constrain the inputs and outputs of AI systems, particularly large language models (LLMs), to block harmful, inaccurate, or off-policy…
Hallucination in generative AI is the production of content that is unsupported, contradicted by an applicable source, factually wrong, internally inconsistent, or otherwise presented without an adequate basis.
HaluEval (Hallucination Evaluation) is a large-scale benchmark for measuring how well large language models (LLMs) can recognize hallucinated content, that is, text that conflicts with a source or cannot be…
HarmBench is a standardized evaluation framework for automated red teaming and robust refusal of large language models (LLMs).
Holden G. Karnofsky is an American philanthropist, nonprofit executive, and writer on artificial intelligence who co-founded the charity evaluator GiveWell (2007) and the grantmaking organization Open…
This article is a defensive and academic survey of how proprietary large language models (LLMs) such as GPT-4, ChatGPT, Claude, and Gemini can be partially copied or have their internal information leaked…
Human-in-the-loop (HITL) describes any arrangement in which a person is a required participant in an automated system's operating cycle rather than a bystander to it.
Humanity's Last Exam (HLE) is a multi-modal AI benchmark of 2,500 public expert-level academic questions (plus a 500-question private holdout, 3,000 in total) spanning more than 100 disciplines
Ian Hogarth is a British technology entrepreneur, investor, and writer who serves as the founding chair of the UK AI Safety Institute
Igor Babuschkin is a German artificial intelligence researcher and engineer who is a co-founder of xAI, the AI company that Elon Musk started in 2023.