Citation and evidence

Harness (AI)

32 min full readUpdated 55 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI AgentsAI BenchmarksDeveloper ToolsModel Evaluation

Cite this article

This article is about software that wraps and evaluates AI models. For the 2015 paper "Explaining and Harnessing Adversarial Examples" and other uses of the word, see the section on other meanings below.

Harness, in artificial intelligence, is the software scaffolding that wraps a machine learning model, most often a large language model, so the model can be put to work for a purpose beyond answering a single prompt. The term covers two closely related but distinct kinds of system. An agent harness, also called an AI harness, is the runtime layer that turns a model into an AI agent: it supplies the loop, the tools, the memory, and the guardrails that let the model take actions in an environment. An evaluation harness, or eval harness, is the framework that administers benchmarks to a model: it formats prompts, runs the model over a task set, extracts and scores the answers, and reports the results. Both borrow their name and their central metaphor from the test harness of classical software engineering, the rig that drives a piece of code under controlled conditions and observes what it does.[9][31]

The common thread is a division of labor: the model supplies the intelligence, and the harness supplies almost everything else needed to use that intelligence or to measure it. Anthropic describes an agent harness as "the software scaffolding around a model: the loop, tools, context management, and guardrails that turn raw intelligence into a working agent."[1] Databricks frames the split as anatomy, calling the model the "brain" and the harness "everything around it that helps the agent operate safely and reliably."[2] A recurring lesson across both senses is that a reported capability number is a property of the model and its harness together, not of the model alone: the same model can look markedly stronger or weaker depending on the harness wrapped around it.[5][2][29]

Etymology and origins

The word "harness" entered English around 1300, meaning personal fighting equipment or the trappings of a war-horse, from Old French. The non-military sense, "fittings for a beast of burden," followed in the early fourteenth century, and the figurative sense of the verb "to harness," glossed by the Online Etymology Dictionary as "to control for use as power," dates from the 1690s.[34] That horse-tack image of channeling a power source is the one every later technical sense reuses.

In software engineering, a test harness is the automation code that exercises a program under test by feeding it inputs and comparing the outputs against expected values. TechTarget describes one as being "made up of test execution engines and test script repositories," while Wikipedia defines it as "a collection of stubs and drivers configured to assist with the testing of an application or component," standing in for production infrastructure that "is either not available or not desired."[32][31] Artificial intelligence borrowed the term in stages. The evaluation sense came first: EleutherAI released its lm-evaluation-harness in 2021, under a title its repository still echoes as "A framework for few-shot evaluation of language models," and "harness" became standard vocabulary for the rig that runs a model over benchmark tasks.[11][12] The agent sense arrived with the autonomous-agent wave of 2023 and 2024, when the wrapper that lets a model act in a loop was more often called a "scaffold." METR, then publishing as ARC Evals, introduced the usage in a July 2023 report: "We call such programs 'scaffolding' and call scaffolding + model combinations 'agents.'"[52] It was writing "agent scaffolding" by December 2023, in a bounty call for agent tasks.[51] The SWE-agent project named its own version the Agent-Computer Interface.[10] "Agent harness" became the mainstream label over 2025 and 2026, as Anthropic described its Claude Agent SDK as "a powerful, general-purpose agent harness" and a wider practitioner literature grew up around designing these systems.[6][9] The concept predates the word: Anthropic's influential December 2024 essay "Building Effective Agents" describes the agent loop in detail without using either "harness" or "scaffold."[4]

Agent harnesses

The agent loop

At the core of every agent harness is a loop. The model is given a task and a set of tools; it responds with an action, such as a tool call; the harness executes that action, captures the result, and returns the result to the model as a new observation; and the cycle repeats until the model signals that it is finished or a stop condition is reached. Anthropic's working definition of an agent is exactly this arrangement: agents "are typically just LLMs using tools based on environmental feedback in a loop."[4] The same essay draws a useful line between workflows, "systems where LLMs and tools are orchestrated through predefined code paths," and agents, "systems where LLMs dynamically direct their own processes and tool usage."[4] A harness can run either, but the agent case is what makes the harness load-bearing, because the model, not the programmer, decides the sequence of steps.

What a harness contains

Beyond the loop, a production agent harness bundles several services around the model. Databricks enumerates eight common building blocks: system prompts; tools and tool execution; sandboxes and execution environments; a filesystem and durable storage; memory and context management; feedback loops and self-verification; guardrails and human-in-the-loop controls; and observability and logging.[2] Context management is repeatedly singled out as the hard part for long-running agents, because the model's context window is finite and a long task can overflow it. Anthropic's Claude Agent SDK, for instance, "has context management capabilities such as compaction, which enables an agent to work on a task without exhausting the context window," a challenge that has made context engineering a discipline in its own right.[6]

Notable agent harnesses

Many widely used coding and agent tools are, structurally, harnesses wrapped around a general-purpose model.

HarnessDeveloperNotes
Claude Code and the Claude Agent SDKAnthropicCommand-line coding agent; Anthropic calls the SDK "a powerful, general-purpose agent harness" [6]
Hermes AgentNous ResearchMIT-licensed agent with a learning loop that "creates skills from experience" and improves them during use; runs in a terminal and through chat platforms [48]
Kilo CodeKiloMIT-licensed agent shipped as a VS Code extension, a JetBrains plugin and a CLI; the CLI is a fork of OpenCode [49]
OpenAI Codex CLIOpenAITerminal-based coding agent
AiderOpen sourceTerminal pair-programming assistant that edits a local git repository
ClineOpen sourceCoding agent that runs inside the VS Code editor
OpenHandsAll Hands AIOpen-source agent platform, formerly OpenDevin
SWE-agentPrinceton NLPResearch harness that introduced the Agent-Computer Interface
PiEarendil (created by Mario Zechner) [41]Minimal open-source TypeScript coding agent whose model gets only four tools (read, write, edit, bash) by default [38]
Oh My PiCan BölükMIT-licensed fork of Pi that bundles language-server, debugger, subagent and browser tools and a hash-anchored edit format [39]
DevinCognitionCommercial autonomous software-engineering agent
Auto-GPTOpen sourceOne of the first popular autonomous agents built on GPT-4 (2023)
OpenClawOpen sourceAgent that connects to messaging apps and acts on the user's behalf, by OpenRouter's description of its listing [42]
DeepSeek HarnessDeepSeekMIT-licensed harness whose tagline is "Everything is a Plugin," released as a developer preview in August 2026 [53]
Command CodeLangbase, Inc., trading as Command CodeCommercial CLI and desktop agent aimed at open-weight models [45]
FreebuffFreebuff, Inc. (formerly Manicode, Inc.; Codebuff is a product name)Advertising-funded free agent (CLI, desktop, web, cloud and chat) built on the Codebuff multi-agent framework [47]

Expanded article table

Harness, scaffold, framework, and orchestrator

The vocabulary here is not fully settled. In loose usage, "harness" and "scaffold" are interchangeable and both mean roughly everything around the model. Hugging Face's agent glossary proposes a finer split: scaffolding is "the behavior-defining layer around the model," meaning the system prompt, the tool descriptions, and how responses are parsed and remembered, while the harness is "the execution layer inside the agent: it calls the model, handles its tool calls, decides when to stop."[3] In that framing, scaffolding is what the model works from and the harness is what runs it. A 2026 analysis by Sanderson Oliveira de Macedo goes further, setting out necessary and sufficient conditions for a system to count as an agent harness: "(i) an agent loop that interleaves reasoning, action, and observation; (ii) a tool interface that lets the model perceive and alter the environment; (iii) context management that decides what enters and leaves the model's window; and (iv) control mechanisms, that is, limits, verification, and deterministic actions, that make the execution more trustworthy, auditable, and contained." The fourth condition is the strictest: a control mechanism counts only if "its effectiveness does not depend on the model choosing to cooperate."[9] The paper then names five things "routinely confused with the agent harness: agent framework, agent SDK, IDE plugin, eval harness, and orchestrator," and rules each out. An agent framework such as CrewAI or LangGraph routes messages between roles without necessarily closing a loop on an external environment; a plain SDK exposes the tool-calling primitive and leaves the loop to the developer; an evaluation harness "acts after the fact, the agent harness during the fact"; and an orchestrator that chains fixed steps fails the adaptive-loop test.[9]

A harness is also distinct from the package that delivers it. An IDE extension is a distribution channel, a plugin installed into an editor, while the harness is the loop and tool layer running inside it, and the same harness is often shipped both ways. OpenRouter's category vocabulary for the apps that use it keeps cli-agent ("Terminal-based coding assistants") and ide-extension ("Editor/IDE integrations") apart, alongside cloud-agent, programming-app and native-app-builder, and several products claim both at once.[43] Cline and Kilo Code are each listed under IDE Extensions and CLI Agents, and Kilo Code's own README describes it as an agent that "meets you everywhere you work: VS Code, JetBrains, and the CLI."[42][49] None of them is a model: each routes its work to a third-party model the user chooses.

Why the harness matters: model versus harness

The single most consequential fact about harnesses is that they change results. Because a benchmark like SWE-bench runs a whole agent rather than a bare model, the reported score reflects the scaffolding as much as the network. As Anthropic's Erik Schluntz put it, "SWE-bench doesn't just evaluate the AI model in isolation, but rather an entire 'agent' system," and "the performance of an agent on SWE-bench can vary significantly based on this scaffolding, even when using the same underlying AI model."[5] Anthropic reached 49% on SWE-bench Verified with a deliberately minimal scaffold, just a prompt, a bash tool, and a file-editing tool, a design meant to hand as much control as possible to the model.[5]

The academic literature makes the point under controlled conditions. The SWE-agent paper holds the language model fixed and varies only the interface between the agent and the computer. With GPT-4 Turbo unchanged, its Agent-Computer Interface resolved 18.00% (54 of 300) of the SWE-bench Lite split and 12.47% (286 of 2,294) of the full test set; on Lite that is "a 6.7-fold improved % Resolved rate" over the same model used through non-interactive retrieval-augmented generation, which scored 2.67%. The paper's headline comparison is on the full benchmark, where its 12.5% is set against a previous best of 3.8%.[10] METR, which evaluates frontier models for dangerous capabilities, treats this dependence as central to its method. To avoid understating what a model can do, and therefore understating its risks, it instructs evaluators that "the model should be provided with the best available scaffolding + tooling, including at a minimum the ability to do chain-of-thought reasoning, use the command line and see the resulting output, and a way to manage long tasks given context-length limits."[22] METR glosses the broader practice, elicitation, as "post-training enhancements to achieve more capable agent performance."[50] Its August 2024 write-up of GPT-4o reports how much of a measured score is scaffolding: of 152 failed runs it classified 78 (51%) as "spurious failures that seem easily fixed with task-agnostic elicitation effort," and patching one such failure mode turned 4 of 10 failed runs into successes.[24] In one study, run on a suite of 195 agency tasks with METR's own basic scaffolding held fixed, OpenAI's post-training of GPT-4 raised the success rate from 5% (plus or minus 2) at the base model to 30% (plus or minus 6) at the June 2023 release, which METR reports as 26 percentage points. METR's own subsequent scaffolding and tooling work added about 8 more points, to 38% (plus or minus 6), an increase it describes as not statistically significant "despite nontrivial effort on our part."[23]

As base models converge in raw capability, this layer matters more rather than less. Databricks observes that "as models converge in raw capability, the harness increasingly determines performance."[2] Harness-Bench, a 2026 study from Peking University and Qiyuan Tech, measures the effect directly. It runs 106 tasks across a full factorial matrix of 6 configurable harnesses and 8 API model backends, 5,088 trajectories in all, plus 106 more for Codex as a model-bound agent. Over the same tasks and the same pool of backends the best harness in the set, NanoBot, scored 76.2 in aggregate against OpenClaw's 52.4, a 23.8-point gap under the same task set and model-backend pool, which the authors read as configuration-level variation; the authors also find that stronger backends show lower variance across harnesses. They conclude that "agent capability should be reported at the model-harness configuration level rather than attributed to the base model alone."[29] A narrower example comes from the edit tool alone: in February 2026 Can Bölük, the developer of Oh My Pi, reported that swapping the file-editing format for one in which every line read back to the model carries "a 2-3 character content hash" that the model then cites when editing, with the model and prompt unchanged, took Grok Code Fast 1 "from 6.7% to 68.3%, a tenfold improvement" on a synthetic bug-fix suite of 180 mutated-source tasks run three times per model. The post tested 16 models, and its own counts do not agree: the title says 15 models improved, the table shows 15 of 16 with a positive delta over patch, and the chart caption reports that the hash-anchored format "beats patch in 14/16 models."[40]

There is a flip side. Because every part of a harness compensates for something the model cannot yet do on its own, a more capable model needs less harness. Anthropic frames harness design as deciding "what belongs in that scaffolding and, as models improve, what you can take out," and observes that "every component in a harness encodes an assumption about what the model can't do on its own," an assumption that grows stale as the model improves.[1][7]

Harness engineering

As agents moved into daily software work, tuning the harness became a discipline of its own. In February 2026 the engineer Mitchell Hashimoto described what he called "harness engineering," "the idea that anytime you find an agent makes a mistake, you take the time to engineer a solution such that the agent never makes that mistake again," for example by adding a tool that lets the agent verify its own work.[8] Anthropic's own experiments show the trade-off at the heart of the practice: a multi-agent harness can be "over 20x more expensive" to run than a single agent while producing output of visibly higher quality.[7]

The relationship also runs in the other direction: instead of engineering the harness around a fixed model, the model can be post-trained to fit a fixed harness. Frontis-MA1, a 35-billion-parameter model with published weights (35B total with about 3B activated, under the noncommercial CC BY-NC 4.0 licence) published on July 31, 2026 by the China-based group FrontisAI (Horizon Research, Frontis.AI, with Tsinghua University), was fine-tuned with execution-grounded supervised and reinforcement learning on the same four program-evolution operators, Draft, Improve, Debug, and Crossover, that its companion OpenMLE-Evo search harness composes into a long-horizon loop at inference time, so that training and harness share one action space.[35][37] The controlled results show how much of the reported capability lives in each layer: on the 22-task Lite split of MLE-bench, holding the OpenMLE-Evo harness fixed and swapping in Frontis-MA1-35B for its Qwen3.6-35B-A3B base model raised the Medal Average, the mean fraction of tasks taking any Kaggle medal, from 39.39% to 60.61%, while holding the model fixed and upgrading the harness to an asynchronous variant with cross-task experience priors added a further 10.6 points, to 71.21%.[35] FrontisAI's model card reports every headline number as a joint property of the model and harness system rather than of the model alone, the same reporting convention Harness-Bench argues for.[36][29]

Measured usage: the OpenRouter app rankings

Harness adoption is hard to measure, because most harnesses are programs that a user installs locally and points at a model provider, reporting nothing to anyone. The one public, continuously updated count comes from OpenRouter, an API gateway that routes requests to many providers and publishes a leaderboard of the applications sending it traffic. The page is headed "App & Agent Rankings" and its list is labelled "Largest public apps and agents opting into usage tracking on OpenRouter."[42]

What the ranking measures, and does not

The unit is tokens, not users, installs, sessions or revenue, and only tokens that pass through OpenRouter. OpenRouter's app-attribution documentation sets out the mechanism: an application identifies itself with an HTTP-Referer header, which the docs call "the primary identifier for rankings" and state is required, because "Without it, no app page will be created and your usage will not appear in rankings."[43] A developer can suppress the listing entirely by sending X-OpenRouter-App-Visibility: hidden, which excludes a newly created app "from the public rankings, the app marketplace, and public app pages" while attribution keeps working for that developer's own analytics.[43] OpenRouter states the position plainly under privacy considerations: "Only public apps, meaning those that send headers, are included in rankings."[43] The board offers daily, weekly and monthly windows.[43]

Three limits follow from that design. First, the sample is self-selected: a harness appears only because its authors chose to send attribution headers and chose not to hide the result. Second, it is a partial view of any single harness, because the same program pointed at a provider's own API, or run on a vendor subscription plan, routes nothing through OpenRouter and contributes nothing to the count. Third, the category labels on each listing are supplied by the app itself through an X-OpenRouter-Categories header drawn from a fixed vocabulary, so they are self-descriptions rather than an editorial taxonomy.[43] A position on this board is a dated snapshot of OpenRouter traffic, not a measure of overall popularity, market share or quality.

The board on 1 October 2026

Read on 1 October 2026, OpenRouter's global app ranking had the same eleven agent harnesses at the top of all three windows, in three different orders. Token totals are as OpenRouter displayed them, where T is trillions and B billions of tokens. OpenRouter states that each window ends with the most recent complete daily bucket and that the date it shows is the newest usage bucket rather than the time the page was rendered, so a later read on the same day returns the same window; figures read on different days differ because the window has advanced.

AppThis weekThis monthToday
Hermes Agent1 (13T)1 (49.3T)1 (2.07T)
Claude Code2 (7.01T)2 (24.3T)3 (1.07T)
Cline3 (6.4T)4 (15.7T)4 (1.05T)
Kilo Code4 (6.16T)3 (16.9T)2 (1.18T)
pi5 (3.36T)5 (9.95T)6 (468B)
omp6 (2.88T)6 (7.95T)7 (423B)
Codex7 (2.57T)7 (6.74T)5 (699B)
Command Code8 (2.2T)11 (2.3T)8 (379B)
Freebuff9 (1.69T)9 (4.91T)10 (255B)
DeepSeek Harness10 (1.65T)10 (4.87T)9 (259B)
OpenClaw11 (1.37T)8 (5.39T)11 (204B)

Expanded article table

Source: OpenRouter app rankings, read 1 October 2026.[42]

The three columns disagree enough to make the point about volatility on their own. Command Code sat eighth over the week with 2.2T tokens but eleventh over the month with 2.3T, meaning nearly all of its monthly volume arrived in that final week; OpenRouter's own Trending panel listed it as the fastest-growing app on the board at ">999%" week over week, against +108% for Cline, +85% for Codex, +73% for Kilo Code, +70% for omp, +53% for pi, +21% for Claude Code and +8% for Hermes Agent.[42] OpenClaw moved the other way, eighth over the month and eleventh over the week. Hermes Agent led every window by a wide margin, with roughly twice Claude Code's monthly tokens.[42]

Vendors do cite these positions. Command Code published a post on 20 July 2026 headed "Command Code is the most used coding agent on OpenRouter," claiming first place in OpenRouter's Programming App, Native App Builders and IDE Extensions categories and second in Coding Agents, CLI Agents, Personal Agents and Productivity.[46] Those are the company's own claims about a board whose ordering shifts weekly; on OpenRouter's reading of 1 October 2026 the app stood eighth overall for the week.[42]

Identifying the entries

Several listings are labelled only with a domain, and the board gives no developer names, so the entries have to be resolved from the products themselves.

omp appears with the domain omp.sh and no other identification. The installer served from that domain is headed "OMP Coding Agent Installer" and sets REPO="can1357/oh-my-pi" and PACKAGE="@oh-my-pi/pi-coding-agent", so the entry is Oh My Pi, Can Bölük's heavily extended fork of Pi, whose own repository describes it as a "Coding agent with the IDE wired in. Built by Stencil Labs."[44][39] The OpenRouter entry therefore sits one row below the project it was forked from.

Command Code appears with the domain commandcode.ai. It is a commercial coding agent distributed on npm as command-code, announced on 25 February 2026 with, in the company's words, "$5M to launch the first coding agent that can continuously learn your coding taste," a capability it attributes to a component it calls taste-1.[45] Its own marketing positions it as a harness rather than a model, and aimed at open-weight ones: one home-page section is headed "harness, not the model," and the product is pitched as "the best coding agent for open models."[54][45]

Freebuff appears with the line "Powerful coding models, funded by ads." It is an advertising-funded free coding agent published under Apache 2.0 by the CodebuffAI organisation on GitHub, spanning a CLI, a desktop app, a browser product, a hosted cloud product and a chat product, and its README states that it "is built on Codebuff, the open multi-agent framework that powers its orchestration, tools, and SDK," and gives the install command npm install -g freebuff.[47]

The remaining entries are established projects: Hermes Agent from Nous Research, Claude Code from Anthropic, Cline, Kilo Code, Codex from OpenAI, Pi from Earendil, DeepSeek Harness from DeepSeek, and OpenClaw. Roo Code, another editor-based coding agent, whose README points readers to Cline as the project "from where Roo Code originated," did not appear in OpenRouter's top twenty in any of the three windows on that date.[55][42]

Evaluation harnesses

An evaluation harness is a framework that measures a model rather than deploying it. It takes a benchmark, which is a dataset paired with a task definition and a scoring rule, turns each example into a prompt, runs the model, extracts an answer from the output, computes a metric, and reports a comparable number.[16] The reason to package this as reusable software is reproducibility: if everyone runs the same benchmark through the same harness, their scores can be compared.

The lm-evaluation-harness

The canonical example is EleutherAI's lm-evaluation-harness, whose repository describes it as "A framework for few-shot evaluation of language models" and whose first archived release, deposited on Zenodo on September 2, 2021, is titled "A framework for few-shot language model evaluation."[11][12] The project says its "original goal was specifically to compare results with" the GPT-3 paper, "Language Models are Few Shot Learners," so that other models could be measured under the same few-shot methodology rather than against numbers copied from a paper; the release is conventionally cited as Gao et al., 2021.[12][11] The project has grown to cover "over 60 standard academic benchmarks for LLMs, with hundreds of subtasks and variants," and it became the backend of Hugging Face's Open LLM Leaderboard, for which it ran the leaderboard's benchmark suites.[12] The harness has been "used in hundreds of papers" and inside dozens of organizations; foundational open-model papers such as Pythia and GPT-NeoX-20B report their benchmark numbers with it.[12][14][15] GPT-NeoX-20B records even the harness version it used, noting that "all evaluations had version 0 in the Evaluation Harness," a level of detail that later work has not consistently matched.[15]

How an evaluation harness works

Under the hood, a harness can score a model in more than one way, and the choice affects the result. For multiple-choice questions it can compare the probabilities the model assigns to each candidate answer, so that no text is generated and the chosen "answer" is simply the option the model finds most likely; the lm-evaluation-harness calls this a log-likelihood request. For open-ended tasks such as grade-school math or code, it instead lets the model generate text until a stop condition, then extracts the final answer and applies a metric like exact match; this is a generate-until request.[12] Seemingly minor decisions at this stage, the exact prompt wording, the number of in-context examples, and whether the answer is read from letter probabilities or from generated text, all move the number.

Other evaluation frameworks

Several other frameworks fill the same niche. OpenAI released Evals in March 2023 alongside GPT-4, "our framework for automated evaluation of AI model performance, to allow anyone to report shortcomings in our models to help guide further improvements"; its repository description calls it "a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks."[20] Its later simple-evals library is deliberately lightweight and, in its own words, "emphasizing the zero-shot, chain-of-thought setting," published "so we can be transparent about the accuracy numbers we're publishing alongside our latest models."[21] That library now carries a deprecation notice dated July 2025 saying it "will no longer be updated for new models or benchmark results," keeping only reference implementations for HealthBench, BrowseComp and SimpleQA.[21] Stanford's HELM, short for Holistic Evaluation of Language Models, is both a large benchmark suite and a runnable evaluation framework, described by the Center for Research on Foundation Models as "an open source Python framework ... for holistic, reproducible and transparent evaluation of foundation models"; its repository notes that "HELM entered maintenance mode on June 1, 2026."[19]

Reproducibility: the same model, different scores

Because so much depends on the harness, the same model run through two different harnesses can post very different scores on what is nominally the same benchmark. The best-documented case involves MMLU, a multiple-choice knowledge test. In June 2023, as the first models were being added to the Open LLM Leaderboard, Hugging Face found that the published LLaMA scores did not match the ones the leaderboard was producing, and traced the gap to three different implementations of MMLU: the original benchmark code, the version in Stanford's HELM, and the version in the lm-evaluation-harness as of January 2023.[13] For the 65-billion-parameter LLaMA model, the three gave 0.636, 0.637, and 0.488, a spread of about fifteen points on the same model answering the same questions.

MMLU implementationLLaMA-7BLLaMA-13BLLaMA-30BLLaMA-65B
Original implementation0.3510.4700.5840.636
HELM0.3390.4710.5830.637
lm-evaluation-harness0.3420.3770.4570.488

Expanded article table

The differences came from choices that sound trivial: whether the prompt includes a topic line, how the answer options are labeled, and, decisively, how the answer is read out. The original code compared "the probabilities predicted by the model, on the four answers only"; HELM had the model generate text and compared that generation with the expected answer, which for MMLU is the answer letter; the harness used "the probabilities of the full answer sequence, with the letter followed by the text of the answer."[13] As Hugging Face concluded, "these different implementations of the same benchmark give widely different numbers and even change the ranking order of the models on the leaderboard!"[13] The episode is the standard cautionary tale that a benchmark score is close to meaningless without naming the harness that produced it.

Agentic evaluation harnesses

For agents, evaluation harnesses have to do more than score text; they have to run the agent's actions and check the state of the world afterward. SWE-bench's harness takes the code patch an agent produces, applies it to a real repository inside a Docker container, and runs the project's test suite; an instance counts as resolved only when every test in the instance's FAIL_TO_PASS list (tests that should change from failing to passing) and its PASS_TO_PASS list (tests that must keep passing) has a pass status.[16][17] The maintainers moved to "a fully containerized evaluation harness using Docker" in June 2024 specifically "for more reproducible evaluations."[18] Other agent benchmarks grade the outcome rather than the transcript in the same way: Terminal-Bench, whose 2.0 release consists of 89 curated tasks, states that its "tests verify that all outcomes described in the instruction have been achieved by testing properties of the final container state; they do not test the agent's commands or console output,"[25] and the tool-use benchmark tau-bench compares a database's final state against an annotated goal, adding "a new metric (pass^k) to evaluate the reliability of agent behavior over multiple trials."[26] The GAIA benchmark for general assistants grades short factual answers against a held-out key, and reported that human respondents answered 92% of its questions correctly against 15% for GPT-4 with plugins.[27] Because the score depends on the agent scaffold that produced the actions and the harness that graded them, an entry on one of these leaderboards is best read as a pairing of a model with a scaffold, not as a property of the model alone.

Standardizing agent evaluation

The dependence of scores on the harness has prompted efforts to standardize it. The Holistic Agent Leaderboard, from a Princeton-led, multi-institution team, offers "a standardized evaluation harness that orchestrates parallel evaluations across hundreds of VMs, reducing evaluation time from weeks to hours while eliminating common implementation bugs," and used it to run "21,730 agent rollouts across 9 models and 9 benchmarks" for "a total cost of about $40,000," arguing that without shared, standardized practice, agent results cannot be compared meaningfully across approaches.[28] The broader recommendation running through this work is to treat the harness as a controlled variable reported alongside the model, rather than an implementation detail hidden behind a single number.[29]

Other meanings and disambiguation

Because the word is common, "harness" turns up in several contexts unrelated to the AI senses above.

The 2015 paper "Explaining and Harnessing Adversarial Examples," by Ian Goodfellow, Jonathon Shlens, and Christian Szegedy, which introduced the Fast Gradient Sign Method, uses "harnessing" as an ordinary verb, meaning to put adversarial examples to work for adversarial training. It has nothing to do with a software harness.[30]

In robotics and manufacturing, a wire harness is a bound bundle of wires and cables; automating its assembly is a well-known open problem, because the material is flexible and every product is slightly different, and it is sometimes cited as a task that robots still struggle to perform. A safety harness is a body-worn assembly for fall protection. Finally, Harness, at harness.io, is a software-delivery company co-founded in 2017 by Jyoti Bansal, who serves as its chief executive; its delivery, testing, security and cost-management platform shares the name by coincidence and is unrelated to the AI concept.[33]

See also

References

  1. ^1 ^2Martin, Lance. "Agent Harness Design: 3 Patterns for Harnessing Claude's Intelligence." Anthropic (Claude Platform blog), April 2, 2026. claude.com/...harnessing-claudes-intelligence
  2. ^1 ^2 ^3 ^4Databricks Staff. "What is an AI Agent Harness?" Databricks, June 17, 2026. databricks.com/...ai-harness
  3. ^Paniego, Sergio, and Aritra Roy Gosthipaty. "Harness, Scaffold, and the AI Agent Terms Worth Getting Right." Hugging Face, May 25, 2026. huggingface.co/...agent-glossary
  4. ^1 ^2 ^3Anthropic. "Building Effective Agents." December 19, 2024. anthropic.com/...building-effective-agents
  5. ^1 ^2 ^3Schluntz, Erik. "Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet." Anthropic, January 6, 2025. anthropic.com/...swe-bench-sonnet
  6. ^1 ^2 ^3Anthropic. "Effective harnesses for long-running agents." November 26, 2025. anthropic.com/...harnesses-for-long-running-agents
  7. ^1 ^2Rajasekaran, Prithvi. "Harness design for long-running application development." Anthropic, March 24, 2026. anthropic.com/...harness-design-long-running-apps
  8. ^Hashimoto, Mitchell. "My AI Adoption Journey." February 5, 2026. mitchellh.com/...my-ai-adoption-journey
  9. ^1 ^2 ^3 ^4de Macedo, Sanderson Oliveira. "What makes a harness a harness: necessary and sufficient conditions for an agent harness." arXiv:2606.10106, June 8, 2026. arxiv.org/...2606.10106
  10. ^1 ^2Yang, John, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering." arXiv:2405.15793, 2024 (NeurIPS 2024). arxiv.org/...2405.15793
  11. ^1 ^2 ^3Gao, Leo, et al. "A framework for few-shot language model evaluation." Zenodo, v0.0.1, September 2, 2021. DOI 10.5281/zenodo.5371629 (10.5281/zenodo.5371628 is the all-versions DOI; the project currently asks to be cited as "The Language Model Evaluation Harness", v0.4.3, July 2024, DOI 10.5281/zenodo.12608602). github.com/...lm-evaluation-harness
  12. ^1 ^2 ^3 ^4 ^5 ^6EleutherAI. "lm-evaluation-harness" (project README). github.com/...lm-evaluation-harness
  13. ^1 ^2 ^3Fourrier, Clémentine, Nathan Habib, Julien Launay, and Thomas Wolf. "What's going on with the Open LLM Leaderboard?" Hugging Face, June 23, 2023. huggingface.co/...open-llm-leaderboard-mmlu
  14. ^Biderman, Stella, et al. "Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling." arXiv:2304.01373, 2023. arxiv.org/...2304.01373
  15. ^1 ^2Black, Sid, et al. "GPT-NeoX-20B: An Open-Source Autoregressive Language Model." arXiv:2204.06745, 2022. arxiv.org/...2204.06745
  16. ^1 ^2Jimenez, Carlos E., John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" ICLR 2024. arXiv:2310.06770. arxiv.org/...2310.06770
  17. ^SWE-bench. "Evaluation" (documentation). swebench.com/...evaluation
  18. ^SWE-bench. Project README (containerized evaluation harness, June 27, 2024 note). github.com/...SWE-bench
  19. ^Liang, Percy, et al. "Holistic Evaluation of Language Models." arXiv:2211.09110, 2022. Stanford Center for Research on Foundation Models. arxiv.org/...2211.09110 ; github.com/...helm
  20. ^OpenAI. "Evals: a framework for evaluating LLMs and LLM systems." March 2023. github.com/...evals
  21. ^1 ^2OpenAI. "simple-evals." github.com/...simple-evals
  22. ^METR. "Guidelines for capability elicitation." March 15, 2024. metr.org/...-guidelines-for-capability-elicitation
  23. ^METR. "Measuring the impact of post-training enhancements." March 15, 2024. metr.org/...-15-measuring-post-impact-enhancements
  24. ^METR. "Details about METR's preliminary evaluation of GPT-4o." August 7, 2024. metr.org/...gpt-4o-report
  25. ^Merrill, Mike A., Alexander G. Shaw, Nicholas Carlini, et al. "Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces." arXiv:2601.11868, January 17, 2026. arxiv.org/...2601.11868
  26. ^Yao, Shunyu, Noah Shinn, Pedram Razavi, and Karthik Narasimhan (Sierra). "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains." arXiv:2406.12045, June 17, 2024. arxiv.org/...2406.12045
  27. ^Mialon, Grégoire, et al. "GAIA: a benchmark for General AI Assistants." arXiv:2311.12983, 2023. arxiv.org/...2311.12983
  28. ^Kapoor, Sayash, Benedikt Stroebl, et al. "Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation." arXiv:2510.11977, October 13, 2025. arxiv.org/...2510.11977
  29. ^1 ^2 ^3 ^4Yao, Yilun, Xinyu Tan, Chao-Hsuan Liu, et al. "Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows." arXiv:2605.27922, May 27, 2026. arxiv.org/...2605.27922
  30. ^Goodfellow, Ian J., Jonathon Shlens, and Christian Szegedy. "Explaining and Harnessing Adversarial Examples." arXiv:1412.6572, 2014 (ICLR 2015). arxiv.org/...1412.6572
  31. ^1 ^2"Test harness." Wikipedia. en.wikipedia.org/...Test_harness
  32. ^Gillis, Alexander S. "What is test harness?" TechTarget. techtarget.com/...test-harness
  33. ^Harness. "Harness: AI for Everything After Code" (company page). harness.io/...about-us
  34. ^"harness (n.)." Online Etymology Dictionary. etymonline.com/...harness
  35. ^1 ^2Yang, J., Jiang, C., et al. "Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering." arXiv:2607.28568, July 30, 2026. arxiv.org/...2607.28568
  36. ^FrontisAI. "Frontis-MA1-35B" (model card). Hugging Face, July 2026. huggingface.co/...Frontis-MA1-35B
  37. ^FrontisAI. "OpenRSI" (repository README). GitHub, July 31, 2026. github.com/...OpenRSI
  38. ^Mario Zechner. "What I learned building an opinionated and minimal coding agent". mariozechner.at, November 30, 2025. mariozechner.at/...2025-11-30-pi-coding-agent
  39. ^1 ^2can1357/oh-my-pi repository README. GitHub. Retrieved September 23, 2026. github.com/...oh-my-pi
  40. ^Bölük, Can. "We improved 15 LLMs at coding in one afternoon. Only the harness changed." Stencil Labs, February 12, 2026. stencil.so/...the-harness-problem (also served at blog.can.ac/...the-harness-problem)
  41. ^Earendil. "Announcing Pi & Lefos." April 8, 2026. earendil.com/...announcing-pi-and-lefos
  42. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8OpenRouter. "App & Agent Rankings." Read October 1, 2026. openrouter.ai/apps
  43. ^1 ^2 ^3 ^4 ^5 ^6OpenRouter. "App Attribution" (documentation). Read October 1, 2026. openrouter.ai/...app-attribution
  44. ^Stencil Labs. omp project site and installer script (`REPO="can1357/oh-my-pi"`). Read October 1, 2026. omp.sh ; omp.sh/install
  45. ^1 ^2 ^3Command Code. "Introducing Command Code: the coding agent that learns your taste." February 25, 2026. commandcode.ai/...ing-agent-that-learns-your-taste
  46. ^Command Code. "Command Code is the most used coding agent on OpenRouter." July 20, 2026. commandcode.ai/...-used-coding-agent-on-openrouter
  47. ^1 ^2CodebuffAI/freebuff repository README. GitHub. Retrieved October 1, 2026. github.com/...freebuff
  48. ^NousResearch/hermes-agent repository README. GitHub. Retrieved October 1, 2026. github.com/...hermes-agent
  49. ^1 ^2Kilo-Org/kilocode repository README. GitHub. Retrieved October 1, 2026. github.com/...kilocode
  50. ^METR. "Autonomy Evaluation Resources." March 15, 2024. metr.org/...2024-03-13-autonomy-evaluation-resources
  51. ^Barnes, Beth, and Megan Kinniment. "Bounty: Diverse hard tasks for LLM agents." METR, December 16, 2023. metr.org/...unty-diverse-hard-tasks-for-llm-agents
  52. ^Barnes, Beth. "New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks." METR, July 31, 2023. metr.org/...2023-08-01-new-report
  53. ^deepseek-ai/deepseek-harness repository. GitHub. Retrieved October 1, 2026. github.com/...deepseek-harness
  54. ^Command Code. Product home page ("harness, not the model"). Read October 1, 2026. commandcode.ai
  55. ^RooCodeInc/Roo-Code repository README. GitHub. Retrieved October 1, 2026. github.com/...Roo-Code

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

9 revisions · v10 · 6,342 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Restamped 2 Oct 2026 after an independent re-read of OpenRouter's rankings corrected the window description and the Freebuff entity name; the page was fact-checked in full on 1 Oct 2026

Cite this page: AI Wiki. "Harness (AI)." aiwiki.ai, updated 1 Oct 2026, fact-checked 1 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/harness

Suggest edit