AI Benchmarks

Explore AI Benchmarks through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Reasoning Models

Articles that also belong to these categories. Counts cover all of AI Benchmarks.

Showing 1-13 of 13 articles

GPQA

GPQA (Graduate-Level Google-Proof Q&A) is a benchmark of expert-written, four-option multiple-choice questions in biology, physics, and chemistry.

Reasoning Models

MATH

MATH is a benchmark of 12,500 competition mathematics problems used to evaluate the mathematical problem-solving ability of machine learning systems, particularly large language models.

Model EvaluationReasoning Models

Mathematical reasoning in AI

Mathematical reasoning in AI is the ability of computer systems to solve mathematical problems: carrying out multi-step calculations, proving theorems, and answering competition or research questions that…

AI ResearchMathematics

MuSR

MuSR (Multistep Soft Reasoning) is a benchmark for evaluating multistep reasoning in large language models, built around long free-text narratives such as murder mysteries, object-placement scenarios, and…

Model EvaluationReasoning Models