Model Evaluation

Explore Model Evaluation through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Code Generation

Articles that also belong to these categories. Counts cover all of Model Evaluation.

Showing 1-6 of 6 articles

Multi-SWE-bench

Multi-SWE-bench is a multilingual benchmark for evaluating the ability of large language model based coding systems to resolve real-world software issues across seven programming languages.

AI BenchmarksAI Code Generation

Pass@k

Pass@k is the standard metric for evaluating code generation models: it measures the probability that at least one of k generated candidate solutions passes all of a problem's unit tests.

AI BenchmarksAI Code Generation

Real-SWE

Real-SWE is a coding-agent benchmark published in September 2026 by Specific Labs, a San Francisco company that licenses operational data and source code from businesses and packages them as training and…

AI BenchmarksAI Code Generation