AI Benchmarks

Explore AI Benchmarks through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Code Generation

Articles that also belong to these categories. Counts cover all of AI Benchmarks.

Showing 1-25 of 25 articles

BigCodeBench

BigCodeBench is a Python code generation benchmark of 1,140 function-level programming tasks that require composing 723 distinct function calls from 139 libraries across seven domains

AI Code Generation

CRMArena / CRMArena-Pro

CRMArena is an AI benchmark for evaluating large language model agents on professional customer relationship management (CRM) tasks inside a realistic, schema-faithful Salesforce environment.

AI Code Generation

CRUXEval

CRUXEval (Code Reasoning, Understanding, and eXecution Evaluation) is a benchmark designed to measure how well large language models can reason about, understand, and mentally execute short Python programs.

AI Code GenerationMachine Learning

FeatureBench

FeatureBench is an execution-based benchmark for measuring how well LLM-powered coding agents handle complex, feature-oriented software development rather than bug fixing.

AI Code Generation

HumanEval

HumanEval is a benchmark for measuring whether a code-generating large language model can complete short Python functions so that they pass unit tests. It contains 164 hand-written tasks.

AI Code Generation

LiveCodeBench

LiveCodeBench is a holistic and contamination-free benchmark for evaluating large language models on code, first released in March 2024 by researchers at UC Berkeley, MIT, and Cornell led by Naman Jain.

AI Code GenerationMachine Learning

MBPP

MBPP (Mostly Basic Python Problems) is a code generation benchmark of 974 crowd-sourced Python programming tasks designed to be solvable by entry-level programmers, introduced by Jacob Austin, Augustus Odena…

AI Code GenerationLarge Language Models

Pass@k

Pass@k is the standard metric for evaluating code generation models: it measures the probability that at least one of k generated candidate solutions passes all of a problem's unit tests.

AI Code GenerationMachine Learning

Real-SWE

Real-SWE is a coding-agent benchmark published in September 2026 by Specific Labs, a San Francisco company that licenses operational data and source code from businesses and packages them as training and…

AI Code GenerationModel Evaluation

RepoBench

RepoBench is an AI benchmark for repository-level code auto-completion, introduced in the 2023 paper "RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems" by Tianyang Liu, Canwen Xu, and…

AI Code Generation

SWE-Atlas

SWE-Atlas is a benchmark for evaluating AI coding agents on professional software-engineering work that goes beyond fixing bugs and resolving issues.

AI AgentsAI Code Generation

SWE-Bench Pro

SWE-Bench Pro (stylized SWE-BENCH PRO) is a contamination-resistant benchmark, released by Scale AI in September 2025, that measures whether an AI coding agent can resolve long-horizon

AI AgentsAI Code Generation

SWE-Lancer

SWE-Lancer is a benchmark released by OpenAI in February 2025 that evaluates the ability of frontier large language models to perform real-world freelance software-engineering work.

AI Code GenerationOpenAI

SWE-rebench

SWE-rebench is a continuously refreshed, contamination-resistant AI benchmark and public leaderboard for evaluating AI agents on real-world software engineering tasks.

AI Code Generation

TheAgentCompany

TheAgentCompany is an AI benchmark that evaluates AI agents on long-horizon, economically valuable knowledge work inside a self-hosted simulation of a small software company.

AI Code Generation

Vibe Code Bench

Vibe Code Bench (VCB) is a benchmark that measures whether AI models can build a complete, deployable web application from nothing but a written specification.

AI Code Generation