AI Benchmarks

Explore AI Benchmarks through related topics and the articles other pages reference most.

Explore articles

Browse subtopics (51)

Articles that also belong to these categories. Counts cover all of AI Benchmarks.

Showing 61-120 of 239 articles

EnigmaEval

EnigmaEval is an AI benchmark of long, complex multimodal puzzles drawn from real-world puzzle hunts, designed to measure the unstructured, creative, multi-step reasoning abilities of frontier AI models.

Model Evaluation

ExploitBench

ExploitBench is a cybersecurity benchmark that measures how far a large language model agent can climb the exploitation "ladder" against a known vulnerability, rather than scoring exploitation as a single pass…

AI SafetyAI in Cybersecurity

FActScore

FActScore (Factual precision in Atomicity Score) is an evaluation method and metric, introduced in 2023, for measuring the factual precision of long-form text generated by large language models.

AI Safety

FLORES-200

FLORES-200 is a multilingual evaluation benchmark for machine translation systems, covering 200 languages across a wide range of language families, scripts, and resource levels.

Natural Language Processing

Factorio Learning Environment

The Factorio Learning Environment (FLE) is an open source benchmark and research framework that uses the industrial automation game Factorio to evaluate the long-horizon agentic capabilities of large language…

FeatureBench

FeatureBench is an execution-based benchmark for measuring how well LLM-powered coding agents handle complex, feature-oriented software development rather than bug fixing.

AI Code Generation

FinanceBench

FinanceBench is an AI benchmark for open-book financial question answering, designed to test whether large language models can answer the kinds of questions a financial analyst asks about a publicly traded…

AI Safety

FormulaOne

FormulaOne is an AI benchmark for evaluating whether a model can design and implement dynamic programming algorithms for graph problems.

Model Evaluation

FrontierMath

FrontierMath is an advanced mathematical reasoning benchmark created by Epoch AI in collaboration with over 60 expert mathematicians, including Fields Medalists Terence Tao, Timothy Gowers, and Richard…

Artificial IntelligenceMathematics

GAIA benchmark

GAIA (General AI Assistants) is a benchmark for evaluating general-purpose AI agents and assistants on real-world tasks that require reasoning, web browsing, file handling, and multimodal understanding.

AI Agents

GDPval

GDPval is a benchmark released by OpenAI on September 25, 2025 that measures how well frontier AI models can produce the actual deliverables of professional knowledge work

AI Research

GPQA

GPQA (Graduate-Level Google-Proof Q&A) is a benchmark of expert-written, four-option multiple-choice questions in biology, physics, and chemistry.

Reasoning Models

GPQA Diamond

GPQA Diamond is the 198-question hardest subset of the Graduate-Level Google-Proof Q&A Benchmark (GPQA), a set of PhD-level multiple-choice questions in biology, physics, and chemistry used to measure…

GSO

GSO (short for Global Software Optimization in the project page subtitle, and Challenging Software Optimization Tasks for Evaluating SWE-Agents in the paper title) is an artificial intelligence benchmark for…

GenAI-Bench

GenAI-Bench is an AI benchmark for evaluating compositional text-to-image and text-to-video generation, introduced in 2024 by researchers from Carnegie Mellon University and Meta AI .

Computer Vision

GeoBench

GeoBench is the umbrella name for a family of artificial intelligence benchmarks evaluating models on geospatial reasoning, earth monitoring, and geographic localization.

HELMET

HELMET (How to Evaluate Long-context Models Effectively and Thoroughly) is a benchmark for evaluating long-context language models introduced by researchers at Princeton University and Intel Labs in 2024.

Model Evaluation

HalluLens

HalluLens is a large language model hallucination benchmark introduced by researchers at Meta AI's Fundamental AI Research (FAIR) lab, together with collaborators at the Hong Kong University of Science and…

Meta AIModel Evaluation

HaluEval

HaluEval (Hallucination Evaluation) is a large-scale benchmark for measuring how well large language models (LLMs) can recognize hallucinated content, that is, text that conflicts with a source or cannot be…

AI SafetyLarge Language Models

HealthBench

HealthBench is an open-source benchmark released by OpenAI on May 12, 2025, that evaluates how large language models handle realistic, multi-turn healthcare conversations.

Healthcare AIOpenAI

HellaSwag

HellaSwag is a commonsense reasoning benchmark for language models, introduced by Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi in the 2019 paper "HellaSwag: Can a Machine Really…

Natural Language Processing

HotpotQA

HotpotQA is a large-scale, multi-hop question answering dataset of about 112,779 crowd-authored question-and-answer pairs over English Wikipedia, whose answers cannot be found in any single paragraph and…

Artificial IntelligenceData & Datasets

HumanEval

HumanEval is a benchmark for measuring whether a code-generating large language model can complete short Python functions so that they pass unit tests. It contains 164 hand-written tasks.

AI Code Generation

Humanity's Last Exam

Humanity's Last Exam (HLE) is a multi-modal AI benchmark of 2,500 public expert-level academic questions (plus a 500-question private holdout, 3,000 in total) spanning more than 100 disciplines

AI Safety

IFBench

IFBench (Instruction Following Benchmark) is an artificial intelligence benchmark that measures whether large language models can follow precise output constraints they have never seen during training.

InferenceX

InferenceX, launched in October 2025 under the name InferenceMAX, is an open-source benchmark that continuously measures large language model inference performance across AI accelerators and serving software.

AI HardwareAI Inference

Iris dataset

The Iris dataset, sometimes referred to as Fisher's Iris dataset or the Iris flower dataset, is a multivariate dataset introduced by the British statistician and biologist Ronald Fisher in his 1936 paper "The…

Data & DatasetsMachine Learning

KernelBench

KernelBench is an AI benchmark and open-source evaluation environment that measures how well large language models can write fast and correct GPU kernels.

Model Evaluation

LAB-Bench

LAB-Bench (the Language Agent Biology Benchmark) is an AI benchmark of more than 2,400 multiple-choice questions built to measure how well language models and AI agents can perform practical biology research…

Model Evaluation

LAMBADA

LAMBADA (LAnguage Modeling Broadened to Account for Discourse Aspects) is a benchmark dataset designed to evaluate the ability of computational language models to understand broad discourse context.

Natural Language Processing

LIBERO

LIBERO ("LIfelong learning BEnchmark on RObot manipulation tasks") is an AI benchmark for studying knowledge transfer in lifelong robot learning.

Robotics

LLM Benchmarks Timeline

The LLM benchmarks timeline is the chronological record of the major tests used to evaluate large language models, from the BERT era reading-comprehension suites of 2018 to the agentic and frontier-reasoning…

LLM Comparisons

LLM comparison is the practice of ranking large language models against one another on quality, cost, speed, and specific skills, using public leaderboards, standardized benchmarks, human preference arenas…

LLM Rankings

LLM rankings order large language models from best to worst using one of two broad methods: human preference voting and fixed benchmark scoring.

LM Evaluation Harness

LM Evaluation Harness (the Language Model Evaluation Harness, often abbreviated lm-eval or written lm-evaluation-harness) is an open-source software framework for measuring the performance of large language…

Model EvaluationOpen Source AI

LMArena

LMArena is a crowdsourced artificial intelligence evaluation platform and company that ranks large language models by having anonymous users vote on which of two blind, side-by-side model responses they prefer

AI CompaniesModel Evaluation