AdvBench
AdvBench (Adversarial Behavior Benchmark) is a red-teaming benchmark dataset for measuring how easily an aligned large language model can be pushed into producing harmful or objectionable content
Explore AI Benchmarks through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Benchmarks.
Showing 1-25 of 25 articles
AdvBench (Adversarial Behavior Benchmark) is a red-teaming benchmark dataset for measuring how easily an aligned large language model can be pushed into producing harmful or objectionable content
Agent benchmark reward hacking refers to the practice of inflating an AI agent's score on an evaluation suite by attacking the evaluation machinery itself rather than by completing the assigned tasks.
AgentDojo is a dynamic evaluation environment for measuring prompt injection attacks and defenses against tool-using large language model agents.
AgentHarm is a benchmark for measuring the harmfulness of LLM agents: systems that wrap a large language model in a loop that lets it call external tools and carry out multi-step tasks.
BBQ (the Bias Benchmark for QA) is a hand-built evaluation dataset that measures whether a question answering (QA) language model relies on social stereotypes when it answers.
Cybench (short for Cybersecurity benchmark) is an open-source evaluation framework for measuring the cybersecurity capabilities and risks of large language model agents.
ExploitBench is a cybersecurity benchmark that measures how far a large language model agent can climb the exploitation "ladder" against a known vulnerability, rather than scoring exploitation as a single pass…
FActScore (Factual precision in Atomicity Score) is an evaluation method and metric, introduced in 2023, for measuring the factual precision of long-form text generated by large language models.
FinanceBench is an AI benchmark for open-book financial question answering, designed to test whether large language models can answer the kinds of questions a financial analyst asks about a publicly traded…
HaluEval (Hallucination Evaluation) is a large-scale benchmark for measuring how well large language models (LLMs) can recognize hallucinated content, that is, text that conflicts with a source or cannot be…
HarmBench is a standardized evaluation framework for automated red teaming and robust refusal of large language models (LLMs).
Humanity's Last Exam (HLE) is a multi-modal AI benchmark of 2,500 public expert-level academic questions (plus a 500-question private holdout, 3,000 in total) spanning more than 100 disciplines
JailbreakBench is an open-source robustness benchmark for evaluating jailbreak attacks and defenses against large language models (LLMs).
LongFact and SAFE are a paired benchmark and evaluation method for measuring the long-form factuality of large language models, introduced by researchers at Google DeepMind and Stanford University in the 2024…
MACHIAVELLI is a benchmark for evaluating the ethical behavior of AI agents in text-based interactive environments.
MASK (Model Alignment between Statements and Knowledge) is an AI safety benchmark that measures the honesty of large language models (LLMs) by testing whether a model will knowingly assert something it…
METR (Model Evaluation and Threat Research) is a nonprofit research organization based in Berkeley, California, that develops scientific methods for measuring the autonomous capabilities of frontier AI systems…
MedHELM (Holistic Evaluation of Large Language Models for Medical Tasks) is a benchmark and evaluation framework that measures how well large language models perform on realistic clinical work.
OpenAI MRCR (Multi-Round Co-reference Resolution) is a long-context evaluation dataset published by OpenAI that measures a language model's ability to distinguish between multiple near-identical "needles"…
SimpleQA is a factuality benchmark released by OpenAI on October 30, 2024 that measures whether large language models can answer short, fact-seeking questions correctly instead of producing hallucinations.
StrongREJECT is a benchmark and automated evaluator for measuring how well jailbreaking attacks actually work against large language models.
The task-completion time horizon is a metric for AI capability proposed by METR that expresses a model's ability in units of human time: it is the length of task
ToxiGen is a large-scale, machine-generated dataset designed for adversarial and implicit hate speech detection.
TruthfulQA is a benchmark designed to measure whether large language models (LLMs) generate truthful answers to questions.
The Weapons of Mass Destruction Proxy (WMDP) is a publicly released benchmark and unlearning testbed for large language models, introduced in the paper "The WMDP Benchmark: Measuring and Reducing Malicious Use…