LM Evaluation Harness
LM Evaluation Harness (the Language Model Evaluation Harness, often abbreviated lm-eval or written lm-evaluation-harness) is an open-source software framework for measuring the performance of large language models on a wide range of benchmarks. Maintained by EleutherAI, it standardizes the steps between a benchmark and a model, formatting each task's prompts, running the model, extracting and scoring the answers, and reporting metrics, so that results are reproducible and comparable across models and across research groups. It became shared infrastructure for open-model research, and it was the backend of Hugging Face's Open LLM Leaderboard until that leaderboard was retired in 2025.[1][4]
The project describes itself as "a framework for few-shot evaluation of language models," and it covers "over 60 standard academic benchmarks for LLMs, with hundreds of subtasks and variants."[1] Its canonical citation is Gao et al., first released in September 2021 and substantially rewritten in December 2023; the README asks to be cited as version 0.4.3 of July 2024. The latest release is version 0.4.13, published on August 31, 2026.[1][2][11][12]
History
EleutherAI, a grassroots research collective formed in 2020, built the harness during its effort to train and release open replications of GPT-3, including the GPT-Neo, GPT-J, and GPT-NeoX models. The group wanted to run its own evaluations rather than copy claimed results from earlier papers, as the GPT-NeoX-20B paper puts it.[6] The first citable release, version 0.0.1, was deposited on Zenodo on September 2, 2021 under the title "A framework for few-shot language model evaluation" and is the "Gao et al., 2021" reference that recurs throughout the open-model literature. Its DOI is 10.5281/zenodo.5371629; the often-quoted 10.5281/zenodo.5371628 is the project's concept DOI, which covers all versions and resolves to whichever release is newest.[12][13]
On December 4, 2023 the project shipped version 0.4.0, a large rewrite that Zenodo titles the "Major refactor," which introduced a task-configuration format based on YAML files and a unified request interface for models.[14] The README's "Cite as" block no longer points at that release: it asks for "The Language Model Evaluation Harness," version 0.4.3 of July 2024, DOI 10.5281/zenodo.12608602, although both the DOI badge at the top of the same README and the repository's CITATION.bib file still point at the v0.4.0 record.[1][2][14] Development has continued steadily since: version 0.4.12 was released on May 11, 2026 and version 0.4.13 on August 31, 2026, a release EleutherAI describes as "a fix-focused release" whose main fixes, for few-shot leakage, a multiple-choice filter bug and group standard error, "may shift previously reported numbers."[11]
How it works
An evaluation run turns each example in a benchmark into a prompt, optionally prepends a fixed number of in-context examples (the few-shot setting), runs the model, and scores the output. The harness asks a model to answer requests of three kinds, and the choice of request determines how a task is measured:[3]
- loglikelihood: given an input and a candidate continuation, return the log-probability the model assigns to that continuation. Multiple-choice tasks use this to rank the candidate answers by probability, with no text generated.
- loglikelihood_rolling: return the log-probability of an entire string, used to compute perplexity over a corpus.
- generate_until: sample text until a stop condition, used for open-ended tasks such as math or code, after which an answer is extracted from the output and scored with a metric such as exact match.
The harness can drive many model backends, among them Hugging Face Transformers models, high-throughput inference servers such as vLLM, and commercial models served behind OpenAI-compatible APIs. To support reproducibility, it versions its task definitions and prompts and can log the exact inputs and outputs for every sample.[1]
Adoption
The harness became shared infrastructure for the open-model community. It is the backend of Hugging Face's Open LLM Leaderboard, which used it to score thousands of open-weight models: first on a 2023 suite of six tasks (ARC, HellaSwag, MMLU, TruthfulQA, Winogrande, and GSM8K), then, after a June 2024 relaunch, on a harder suite (IFEval, BBH, MATH, GPQA, MuSR, and MMLU-Pro) before the leaderboard stopped taking submissions in 2025.[4][7] By EleutherAI's own account the harness "has been used in hundreds of papers" and "is used internally by dozens of organizations including NVIDIA, Cohere, BigScience, BigCode, Nous Research, and Mosaic ML."[1] Foundational open-model papers, including Pythia and GPT-NeoX-20B, report their benchmark numbers with it.[5][6]
Reproducibility and why the harness matters
Because a benchmark score depends on details that the harness fixes, the prompt wording, the number of shots, and above all how the answer is read out, the same model can post different numbers under different harnesses. The best-known illustration came in June 2023, when Hugging Face traced a discrepancy in reported MMLU scores to three different implementations of that benchmark: the original code, HELM, and the LM Evaluation Harness each scored the same LLaMA model differently, in one case by about fifteen points, because they read the model's answer in different ways.[4] The episode is a standard argument that a benchmark result is only meaningful when the harness that produced it is named. The community later adjusted the harness's MMLU implementation to match the original one: as Hugging Face put it, contributors "have done an amazing work updating the evaluation of MMLU in the harness to make it similar to the original implementation and match these numbers."[4]
Related tools
Several other frameworks occupy the same space. Stanford's HELM is a comparable framework that emphasizes reporting many metrics at once. OpenAI's Evals is an open framework with a registry of evals for testing OpenAI models and writing custom ones, and Hugging Face's lighteval, influenced by both the LM Evaluation Harness and HELM, helped run the second version of the Open LLM Leaderboard.[8][9][10] These frameworks differ in their default methodology, for example few-shot log-likelihood scoring versus zero-shot chain-of-thought generation, which is one reason their scores are not directly interchangeable.
See also
- Harness (AI)
- Benchmark (AI)
- Model Evaluation
- EleutherAI
- MMLU
- HELM (Holistic Evaluation of Language Models)
References
- ^1 ^2 ^3 ^4 ^5 ^6EleutherAI. "lm-evaluation-harness" (project README). GitHub. github.com/...lm-evaluation-harness
- ^1 ^2Gao, Leo, et al. "The Language Model Evaluation Harness." Zenodo, version 0.4.3, July 2024. DOI 10.5281/zenodo.12608602. The citation requested in the project README. zenodo.org/...12608602
- ^EleutherAI. "Model guide" (request types: loglikelihood, loglikelihood_rolling, generate_until). lm-evaluation-harness documentation. github.com/...model_guide.md
- ^1 ^2 ^3 ^4Fourrier, Clémentine, Nathan Habib, Julien Launay, and Thomas Wolf. "What's going on with the Open LLM Leaderboard?" Hugging Face, June 23, 2023. huggingface.co/...open-llm-leaderboard-mmlu
- ^Biderman, Stella, et al. "Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling." arXiv:2304.01373, 2023. arxiv.org/...2304.01373
- ^1 ^2Black, Sid, et al. "GPT-NeoX-20B: An Open-Source Autoregressive Language Model." arXiv:2204.06745, 2022. arxiv.org/...2204.06745
- ^Hugging Face. "Open LLM Leaderboard" (about and archive documentation). huggingface.co/...about
- ^Liang, Percy, et al. "Holistic Evaluation of Language Models." arXiv:2211.09110, 2022. Stanford Center for Research on Foundation Models. arxiv.org/...2211.09110
- ^OpenAI. "Evals: a framework for evaluating LLMs and LLM systems." github.com/...evals
- ^Hugging Face. "lighteval." github.com/...lighteval
- ^1 ^2EleutherAI. "Releases" (v0.4.13 published August 31, 2026, with release notes; v0.4.12 published May 11, 2026). lm-evaluation-harness, GitHub. Retrieved October 1, 2026. github.com/...releases
- ^1 ^2Gao, Leo, et al. "A framework for few-shot language model evaluation." Zenodo, version 0.0.1, issued September 2, 2021. DOI 10.5281/zenodo.5371629. doi.org/...zenodo.5371629
- ^"EleutherAI/lm-evaluation-harness." Zenodo concept (all-versions) DOI 10.5281/zenodo.5371628, which resolves to the most recent release (v0.4.13, issued August 31, 2026, as of October 1, 2026). DataCite record retrieved October 1, 2026. doi.org/...zenodo.5371628
- ^1 ^2"EleutherAI/lm-evaluation-harness: Major refactor." Zenodo, version 0.4.0, issued December 4, 2023. DOI 10.5281/zenodo.10256836. doi.org/...zenodo.10256836
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
4 revisions · v5 · 1,235 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Full-page independent fact-check 2 Oct 2026 against the repo README, the GitHub API and DataCite; verdict ok, 6 presentational defects corrected incl. the retired Open LLM Leaderboard backend claim
Cite this page: AI Wiki. "LM Evaluation Harness." aiwiki.ai, updated 1 Oct 2026, fact-checked 1 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/lm_evaluation_harness