AI Benchmarks

Explore AI Benchmarks through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Safety

Articles that also belong to these categories. Counts cover all of AI Benchmarks.

Showing 1-25 of 25 articles

AdvBench

AdvBench (Adversarial Behavior Benchmark) is a red-teaming benchmark dataset for measuring how easily an aligned large language model can be pushed into producing harmful or objectionable content

AI SafetyLarge Language Models

Agent benchmark reward hacking

Agent benchmark reward hacking refers to the practice of inflating an AI agent's score on an evaluation suite by attacking the evaluation machinery itself rather than by completing the assigned tasks.

AI AgentsAI Safety

AgentDojo

AgentDojo is a dynamic evaluation environment for measuring prompt injection attacks and defenses against tool-using large language model agents.

AI AgentsAI Safety

AgentHarm

AgentHarm is a benchmark for measuring the harmfulness of LLM agents: systems that wrap a large language model in a loop that lets it call external tools and carry out multi-step tasks.

AI AgentsAI Safety

Cybench

Cybench (short for Cybersecurity benchmark) is an open-source evaluation framework for measuring the cybersecurity capabilities and risks of large language model agents.

AI SafetyModel Evaluation

ExploitBench

ExploitBench is a cybersecurity benchmark that measures how far a large language model agent can climb the exploitation "ladder" against a known vulnerability, rather than scoring exploitation as a single pass…

AI SafetyAI in Cybersecurity

FActScore

FActScore (Factual precision in Atomicity Score) is an evaluation method and metric, introduced in 2023, for measuring the factual precision of long-form text generated by large language models.

AI Safety

FinanceBench

FinanceBench is an AI benchmark for open-book financial question answering, designed to test whether large language models can answer the kinds of questions a financial analyst asks about a publicly traded…

AI Safety

HaluEval

HaluEval (Hallucination Evaluation) is a large-scale benchmark for measuring how well large language models (LLMs) can recognize hallucinated content, that is, text that conflicts with a source or cannot be…

AI SafetyLarge Language Models

Humanity's Last Exam

Humanity's Last Exam (HLE) is a multi-modal AI benchmark of 2,500 public expert-level academic questions (plus a 500-question private holdout, 3,000 in total) spanning more than 100 disciplines

AI Safety

LongFact / SAFE

LongFact and SAFE are a paired benchmark and evaluation method for measuring the long-form factuality of large language models, introduced by researchers at Google DeepMind and Stanford University in the 2024…

AI Safety

MASK

MASK (Model Alignment between Statements and Knowledge) is an AI safety benchmark that measures the honesty of large language models (LLMs) by testing whether a model will knowingly assert something it…

AI Safety

METR

METR (Model Evaluation and Threat Research) is a nonprofit research organization based in Berkeley, California, that develops scientific methods for measuring the autonomous capabilities of frontier AI systems…

AI SafetyResearch Organizations

MedHELM

MedHELM (Holistic Evaluation of Large Language Models for Medical Tasks) is a benchmark and evaluation framework that measures how well large language models perform on realistic clinical work.

AI Safety

SimpleQA

SimpleQA is a factuality benchmark released by OpenAI on October 30, 2024 that measures whether large language models can answer short, fact-seeking questions correctly instead of producing hallucinations.

AI SafetyNatural Language Processing

ToxiGen

ToxiGen is a large-scale, machine-generated dataset designed for adversarial and implicit hate speech detection.

AI EthicsAI Safety

WMDP benchmark

The Weapons of Mass Destruction Proxy (WMDP) is a publicly released benchmark and unlearning testbed for large language models, introduced in the paper "The WMDP Benchmark: Measuring and Reducing Malicious Use…

AI Safety