AI Benchmarks

Explore AI Benchmarks through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Mathematics

Articles that also belong to these categories. Counts cover all of AI Benchmarks.

Showing 1-4 of 4 articles

FrontierMath

FrontierMath is an advanced mathematical reasoning benchmark created by Epoch AI in collaboration with over 60 expert mathematicians, including Fields Medalists Terence Tao, Timothy Gowers, and Richard…

Artificial IntelligenceMathematics

MATH-500

MATH-500 is a 500-problem benchmark for evaluating the mathematical reasoning of large language models, formed by holding out 500 problems from the test split of the MATH benchmark of Dan Hendrycks et al.

Mathematics

Mathematical reasoning in AI

Mathematical reasoning in AI is the ability of computer systems to solve mathematical problems: carrying out multi-step calculations, proving theorems, and answering competition or research questions that…

AI ResearchMathematics