AI Benchmarks

Explore AI Benchmarks through related topics and the articles other pages reference most.

Most referenced in this topic

Ranked by links from other AI Wiki pages.

Explore articles

Browse subtopics (51)

Articles that also belong to these categories. Counts cover all of AI Benchmarks.

Showing 1-60 of 239 articles

AA-LCR

AA-LCR (Artificial Analysis Long Context Reasoning) is a benchmark for large language models that evaluates the ability to reason across multiple real-world documents totalling approximately 100,000 tokens per…

Natural Language Processing

AGIEval

AGIEval is an AI benchmark for evaluating foundation models on tasks that were originally designed for, and taken by, humans.

Model Evaluation

AIME 2024

AIME 2024 is an AI benchmark of 30 problems drawn from the 2024 American Invitational Mathematics Examination that has become the standard yardstick for measuring the mathematical reasoning of large language…

AIME 2025

AIME 2025 is a 30-problem mathematical reasoning AI benchmark built from the 2025 American Invitational Mathematics Examination

ARC-AGI

ARC-AGI (Abstraction and Reasoning Corpus for Artificial General Intelligence) is a family of AI benchmarks, created by Francois Chollet, that measures fluid intelligence: the ability to solve genuinely novel…

Artificial Intelligence

ARC-AGI 3

ARC-AGI 3 is an interactive reasoning benchmark published by the ARC Prize Foundation and designed to measure how efficiently an artificial system can acquire new skills inside novel, turn-based game…

AdvBench

AdvBench (Adversarial Behavior Benchmark) is a red-teaming benchmark dataset for measuring how easily an aligned large language model can be pushed into producing harmful or objectionable content

AI SafetyLarge Language Models

Agent benchmark reward hacking

Agent benchmark reward hacking refers to the practice of inflating an AI agent's score on an evaluation suite by attacking the evaluation machinery itself rather than by completing the assigned tasks.

AI AgentsAI Safety

Agent evaluation

Agent evaluation is the systematic measurement of how well AI agents (LLM-based systems that plan and act over multiple steps using tools) perform on real-world tasks, using benchmarks, metrics, and testing…

AI AgentsModel Evaluation

AgentDojo

AgentDojo is a dynamic evaluation environment for measuring prompt injection attacks and defenses against tool-using large language model agents.

AI AgentsAI Safety

AgentHarm

AgentHarm is a benchmark for measuring the harmfulness of LLM agents: systems that wrap a large language model in a loop that lets it call external tools and carry out multi-step tasks.

AI AgentsAI Safety

Aider Polyglot

Aider Polyglot is a coding benchmark that evaluates large language models on their ability to write and edit code across six programming languages: C++, Go, Java, JavaScript, Python, and Rust.

Arena-Hard

Arena-Hard (and its evaluation tool Arena-Hard-Auto) is an automatic large language model (LLM) benchmark developed by the team behind Chatbot Arena that scores instruction-tuned models on 500 challenging

Model Evaluation

BALROG

BALROG (Benchmarking Agentic LLM and VLM Reasoning On Games) is a benchmark suite for evaluating the agentic capabilities of large language models and vision language models inside long-horizon

BELEBELE

Belebele is a multiple-choice machine reading comprehension (MRC) AI benchmark that is fully parallel across 122 language variants, meaning the same questions, passages, and answer choices are translated into…

Computer Vision

BIG-Bench

BIG-Bench (Beyond the Imitation Game Benchmark) is a large-scale, collaborative benchmark of 204 tasks, contributed by 450 authors across 132 institutions, built to measure and extrapolate the capabilities of…

Large Language ModelsMachine Learning

BLINK

BLINK is an AI benchmark that evaluates the core visual perception abilities of multimodal large language models (MLLMs).

Computer Vision

Benchmark (AI)

In artificial intelligence and machine learning, a benchmark is a specified evaluation used to compare systems under common conditions.

Model Evaluation

BigCodeBench

BigCodeBench is a Python code generation benchmark of 1,140 function-level programming tasks that require composing 723 distinct function calls from 139 libraries across seven domains

AI Code Generation

BigFinanceBench

BigFinanceBench (BFB) is a benchmark for evaluating how well AI agents perform the research work of professional financial analysts, developed by the finance-AI company Rogo and announced on May 27, 2026.

Finance AIModel Evaluation

BoolQ

BoolQ (Boolean Questions) is a natural language processing benchmark dataset of 15,942 naturally occurring yes/no question answering examples, each pairing a real Google search query with a Wikipedia passage…

Natural Language Processing

BountyBench

BountyBench is a cybersecurity benchmark from Stanford University that measures the offensive and defensive capabilities of AI agents on real-world bug-bounty tasks

AI Agents

BrowseComp

BrowseComp (short for Browsing Competition) is a benchmark for measuring how well AI agents can navigate the open internet to retrieve hard-to-find facts.

OpenAI

BrowserGym

BrowserGym is an open-source Gymnasium-style environment and unified benchmark ecosystem for web agent research, developed by ServiceNow Research.

AI Agents

CIFAR-10

CIFAR-10 is a labeled dataset of 60,000 small color images sorted into 10 mutually exclusive object categories, with 6,000 images per class, used as a standard benchmark for image classification.

Computer VisionData & Datasets

CLIP Score

CLIP Score (also written CLIPScore or CLIP-S) is a reference-free automatic evaluation metric that measures how well a text caption matches an image, computed as the rescaled cosine similarity of the image and…

Computer VisionImage Generation

COLLIE

COLLIE (Systematic Construction of Constrained Text Generation Tasks) is a grammar-based benchmark framework for evaluating how well large language models can produce text that satisfies rich

CRMArena / CRMArena-Pro

CRMArena is an AI benchmark for evaluating large language model agents on professional customer relationship management (CRM) tasks inside a realistic, schema-faithful Salesforce environment.

AI Code Generation

CRUXEval

CRUXEval (Code Reasoning, Understanding, and eXecution Evaluation) is a benchmark designed to measure how well large language models can reason about, understand, and mentally execute short Python programs.

AI Code GenerationMachine Learning

CaliBench

CaliBench is a benchmark for image-to-video generative models that asks whether a model reproduces the correct distribution of physical outcomes across many generations from the same starting frame

Model EvaluationVideo Generation

CharXiv

CharXiv is a benchmark for evaluating chart understanding in multimodal large language models (MLLMs), built by researchers at Princeton Language and Intelligence with collaborators at the University of…

Data & Datasets

Chatbot Arena

Chatbot Arena (now branded simply as Arena, and previously known as LMArena) is a crowdsourced evaluation platform for large language models that ranks AI systems based on human preferences through anonymous…

Large Language Models

ChemBench

ChemBench is an automated AI benchmark that measures the chemical knowledge, reasoning, and safety judgment of large language models and compares their performance against expert human chemists.

Model Evaluation

CommonsenseQA

CommonsenseQA is a multiple-choice question answering benchmark of 12,247 questions, introduced in 2019 by Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant

Natural Language Processing

Creative Writing v3

Creative Writing v3 is an artificial intelligence benchmark that evaluates creative writing in large language models (LLMs) using a hybrid framework combining isolated rubric scoring with pairwise Elo…

Cybench

Cybench (short for Cybersecurity benchmark) is an open-source evaluation framework for measuring the cybersecurity capabilities and risks of large language model agents.

AI SafetyModel Evaluation

Deep Research Bench

Deep Research Bench (DRB) is a benchmark introduced by FutureSearch in May 2025 to evaluate how well large language model agents complete complex research tasks on the open web.

DeepResearch Bench

DeepResearch Bench is a benchmark for evaluating long-form reports produced by web research agents.

DesignArena

DesignArena (also written Design Arena) is a crowdsourced benchmark and consumer creation platform for AI-generated design, built by the San Francisco startup Intelligence.

AI CompaniesModel Evaluation

Dynabench

Dynabench is an open-source artificial intelligence benchmarking platform that runs in a web browser and supports human-and-model-in-the-loop dataset creation

EQ-Bench 3

EQ-Bench 3 is an artificial intelligence benchmark that measures the emotional intelligence of large language models through challenging multi-turn role-plays and transcript-analysis tasks

ERQA

ERQA (Embodied Reasoning Question Answering) is a multimodal benchmark released by Google DeepMind in March 2025 to evaluate the embodied reasoning capabilities of vision-language models (VLMs) on robotics…

Embodied AIGoogle DeepMind

EgoSchema

EgoSchema is a diagnostic benchmark for evaluating very long-form video language understanding, introduced by Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik at UC Berkeley.

Computer VisionMultimodal AI