AI Benchmarks

Explore AI Benchmarks through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Agents

Articles that also belong to these categories. Counts cover all of AI Benchmarks.

Showing 1-28 of 28 articles

Agent benchmark reward hacking

Agent benchmark reward hacking refers to the practice of inflating an AI agent's score on an evaluation suite by attacking the evaluation machinery itself rather than by completing the assigned tasks.

AI AgentsAI Safety

Agent evaluation

Agent evaluation is the systematic measurement of how well AI agents (LLM-based systems that plan and act over multiple steps using tools) perform on real-world tasks, using benchmarks, metrics, and testing…

AI AgentsModel Evaluation

AgentDojo

AgentDojo is a dynamic evaluation environment for measuring prompt injection attacks and defenses against tool-using large language model agents.

AI AgentsAI Safety

AgentHarm

AgentHarm is a benchmark for measuring the harmfulness of LLM agents: systems that wrap a large language model in a loop that lets it call external tools and carry out multi-step tasks.

AI AgentsAI Safety

BountyBench

BountyBench is a cybersecurity benchmark from Stanford University that measures the offensive and defensive capabilities of AI agents on real-world bug-bounty tasks

AI Agents

BrowserGym

BrowserGym is an open-source Gymnasium-style environment and unified benchmark ecosystem for web agent research, developed by ServiceNow Research.

AI Agents

GAIA benchmark

GAIA (General AI Assistants) is a benchmark for evaluating general-purpose AI agents and assistants on real-world tasks that require reasoning, web browsing, file handling, and multimodal understanding.

AI Agents

MLE-bench

MLE-bench is a benchmark for measuring how well autonomous AI agents perform real-world machine learning engineering, released by OpenAI's Preparedness team on October 9, 2024 (arXiv:2410.07095).

AI AgentsOpenAI

MM-BrowseComp

MM-BrowseComp is a benchmark for evaluating multimodal web-browsing AI agents, introduced in August 2025 by researchers from ByteDance, Nanjing University, M-A-P, the Institute of Automation of the Chinese…

AI AgentsMultimodal AI

Mind2Web

Mind2Web is the first large-scale dataset and benchmark for building and evaluating generalist AI agents that follow natural language instructions to complete tasks on real websites.

AI Agents

OSWorld

OSWorld is a benchmark for evaluating multimodal AI agents on open-ended tasks performed inside real computer environments.

AI Agents

Online-Mind2Web

Online-Mind2Web is a live-web benchmark for measuring whether browser-based AI agents can complete multi-step tasks on real websites.

AI Agents

PaperBench

PaperBench is a benchmark released by OpenAI in April 2025 that measures whether AI agents can replicate cutting-edge machine learning research papers from scratch.

AI AgentsOpenAI

RSI-Exam

RSI-Exam is a benchmark that tests whether an AI agent can take a working but weak executable method, improve it through hours of autonomous experimentation, and produce a version that still performs better on…

AI Agents

SWE-Atlas

SWE-Atlas is a benchmark for evaluating AI coding agents on professional software-engineering work that goes beyond fixing bugs and resolving issues.

AI AgentsAI Code Generation

SWE-Bench Pro

SWE-Bench Pro (stylized SWE-BENCH PRO) is a contamination-resistant benchmark, released by Scale AI in September 2025, that measures whether an AI coding agent can resolve long-horizon

AI AgentsAI Code Generation

SkillsBench

SkillsBench is a benchmark for measuring whether Agent Skills, the structured packages of procedural knowledge that augment AI agents at inference time, actually improve how well those agents do real work.

AI Agents

Tau2-bench

τ²-bench (also written Tau2-bench or τ^2-bench) is a benchmark for evaluating conversational AI agents in dual-control environments, where both the agent and a simulated user can call tools to read from and…

AI Agents

Vending-Bench

Vending-Bench is a benchmark that measures the long-term coherence of large language model-based AI agents by tasking them with running a simulated vending-machine business over an extended horizon.

AI Agents

VisualWebArena

VisualWebArena (often abbreviated VWA) is a benchmark of 910 realistic, visually grounded web tasks for evaluating multimodal autonomous AI agents, released in January 2024 by researchers at Carnegie Mellon…

AI Agents

WebArena

WebArena is a realistic, self-hosted web environment and benchmark designed for developing and evaluating autonomous AI agents that perform tasks on the web.

AI AgentsModel Evaluation

WebVoyager

WebVoyager is an end-to-end web agent and its companion benchmark, introduced by Hongliang He and seven coauthors in a paper accepted to ACL 2024 .

AI AgentsModel Evaluation

Windows Agent Arena

Windows Agent Arena (WAA) is a reproducible benchmark and evaluation environment created by Microsoft for testing multimodal computer-use agents on real Windows tasks at scale.

AI AgentsMicrosoft

WindowsWorld

WindowsWorld is a process-aware benchmark for evaluating autonomous computer-use agents on professional workflows that cross several Windows applications.

AI Agents