BrowseComp
BrowseComp (short for Browsing Competition) is a benchmark for measuring how well AI agents can navigate the open internet to retrieve hard-to-find facts.
Explore AI Benchmarks through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Benchmarks.
Showing 1-6 of 6 articles
BrowseComp (short for Browsing Competition) is a benchmark for measuring how well AI agents can navigate the open internet to retrieve hard-to-find facts.
HealthBench is an open-source benchmark released by OpenAI on May 12, 2025, that evaluates how large language models handle realistic, multi-turn healthcare conversations.
HealthBench Hard is a 1,000 example curated subset of the HealthBench benchmark released by OpenAI on May 12, 2025.
MLE-bench is a benchmark for measuring how well autonomous AI agents perform real-world machine learning engineering, released by OpenAI's Preparedness team on October 9, 2024 (arXiv:2410.07095).
PaperBench is a benchmark released by OpenAI in April 2025 that measures whether AI agents can replicate cutting-edge machine learning research papers from scratch.
SWE-Lancer is a benchmark released by OpenAI in February 2025 that evaluates the ability of frontier large language models to perform real-world freelance software-engineering work.