AI Benchmarks

Explore AI Benchmarks through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: OpenAI

Articles that also belong to these categories. Counts cover all of AI Benchmarks.

Showing 1-6 of 6 articles

BrowseComp

BrowseComp (short for Browsing Competition) is a benchmark for measuring how well AI agents can navigate the open internet to retrieve hard-to-find facts.

OpenAI

HealthBench

HealthBench is an open-source benchmark released by OpenAI on May 12, 2025, that evaluates how large language models handle realistic, multi-turn healthcare conversations.

Healthcare AIOpenAI

MLE-bench

MLE-bench is a benchmark for measuring how well autonomous AI agents perform real-world machine learning engineering, released by OpenAI's Preparedness team on October 9, 2024 (arXiv:2410.07095).

AI AgentsOpenAI

PaperBench

PaperBench is a benchmark released by OpenAI in April 2025 that measures whether AI agents can replicate cutting-edge machine learning research papers from scratch.

AI AgentsOpenAI

SWE-Lancer

SWE-Lancer is a benchmark released by OpenAI in February 2025 that evaluates the ability of frontier large language models to perform real-world freelance software-engineering work.

AI Code GenerationOpenAI