AI Agents

Explore AI Agents through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Benchmarks

Articles that also belong to these categories. Counts cover all of AI Agents.

Showing 1-28 of 28 articles

Agent evaluation

Agent evaluation is the systematic measurement of how well AI agents (LLM-based systems that plan and act over multiple steps using tools) perform on real-world tasks, using benchmarks, metrics, and testing…

AI BenchmarksModel Evaluation

AgentDojo

AgentDojo is a dynamic evaluation environment for measuring prompt injection attacks and defenses against tool-using large language model agents.

AI BenchmarksAI Safety

AgentHarm

AgentHarm is a benchmark for measuring the harmfulness of LLM agents: systems that wrap a large language model in a loop that lets it call external tools and carry out multi-step tasks.

AI BenchmarksAI Safety

BountyBench

BountyBench is a cybersecurity benchmark from Stanford University that measures the offensive and defensive capabilities of AI agents on real-world bug-bounty tasks

AI Benchmarks

BrowserGym

BrowserGym is an open-source Gymnasium-style environment and unified benchmark ecosystem for web agent research, developed by ServiceNow Research.

AI Benchmarks

GAIA benchmark

GAIA (General AI Assistants) is a benchmark for evaluating general-purpose AI agents and assistants on real-world tasks that require reasoning, web browsing, file handling, and multimodal understanding.

AI Benchmarks

MLE-bench

MLE-bench is a benchmark for measuring how well autonomous AI agents perform real-world machine learning engineering, released by OpenAI's Preparedness team on October 9, 2024 (arXiv:2410.07095).

AI BenchmarksOpenAI

MM-BrowseComp

MM-BrowseComp is a benchmark for evaluating multimodal web-browsing AI agents, introduced in August 2025 by researchers from ByteDance, Nanjing University, M-A-P, the Institute of Automation of the Chinese…

AI BenchmarksMultimodal AI

Mind2Web

Mind2Web is the first large-scale dataset and benchmark for building and evaluating generalist AI agents that follow natural language instructions to complete tasks on real websites.

AI Benchmarks

OSWorld

OSWorld is a benchmark for evaluating multimodal AI agents on open-ended tasks performed inside real computer environments.

AI Benchmarks

Online-Mind2Web

Online-Mind2Web is a live-web benchmark for measuring whether browser-based AI agents can complete multi-step tasks on real websites.

AI Benchmarks

PaperBench

PaperBench is a benchmark released by OpenAI in April 2025 that measures whether AI agents can replicate cutting-edge machine learning research papers from scratch.

AI BenchmarksOpenAI

RSI-Exam

RSI-Exam is a benchmark that tests whether an AI agent can take a working but weak executable method, improve it through hours of autonomous experimentation, and produce a version that still performs better on…

AI Benchmarks

SWE-Bench Pro

SWE-Bench Pro (stylized SWE-BENCH PRO) is a contamination-resistant benchmark, released by Scale AI in September 2025, that measures whether an AI coding agent can resolve long-horizon

AI BenchmarksAI Code Generation

SkillsBench

SkillsBench is a benchmark for measuring whether Agent Skills, the structured packages of procedural knowledge that augment AI agents at inference time, actually improve how well those agents do real work.

AI Benchmarks

Tau2-bench

τ²-bench (also written Tau2-bench or τ^2-bench) is a benchmark for evaluating conversational AI agents in dual-control environments, where both the agent and a simulated user can call tools to read from and…

AI Benchmarks

Vending-Bench

Vending-Bench is a benchmark that measures the long-term coherence of large language model-based AI agents by tasking them with running a simulated vending-machine business over an extended horizon.

AI Benchmarks

VisualWebArena

VisualWebArena (often abbreviated VWA) is a benchmark of 910 realistic, visually grounded web tasks for evaluating multimodal autonomous AI agents, released in January 2024 by researchers at Carnegie Mellon…

AI Benchmarks

WebArena

WebArena is a realistic, self-hosted web environment and benchmark designed for developing and evaluating autonomous AI agents that perform tasks on the web.

AI BenchmarksModel Evaluation

Windows Agent Arena

Windows Agent Arena (WAA) is a reproducible benchmark and evaluation environment created by Microsoft for testing multimodal computer-use agents on real Windows tasks at scale.

AI BenchmarksMicrosoft

WindowsWorld

WindowsWorld is a process-aware benchmark for evaluating autonomous computer-use agents on professional workflows that cross several Windows applications.

AI Benchmarks