Revision History
Harness (AI) · 10 revisions
Sizes are character counts of the article source. The signed number is the change from the previous revision.
Recent edit summaries
Detailed summaries recorded by editors. Generic maintenance summaries are omitted here; every recorded revision remains below. Dates describe the edit, not necessarily the event it covers.
- Correction: OpenRouter's trailing windows end at the most recent complete daily bucket rather than aggregating through the day; the Freebuff entity went from Manicode, Inc. to Freebuff, Inc. directly
Version 10 · Oct 1, 2026, 06:52 PM
- Update: link the new Command Code and Freebuff articles; the Freebuff vendor is now Freebuff, Inc.
Version 9 · Oct 1, 2026, 06:10 PM
- Correction: Frontis-MA1 is CC BY-NC 4.0 rather than open-weight; the Harness-Bench gap is configuration-level variation, not harness-only; remove unsupported claims and complete a truncated quote
Version 8 · Oct 1, 2026, 01:18 PM
- Correction: note that OpenRouter's trailing-window token totals keep aggregating during the day, so they are an intraday snapshot
Version 7 · Oct 1, 2026, 01:01 PM
- Correction: fixed reference 1's title, the METR scaffolding chronology (July/December 2023, not early 2024), the de Macedo paper's boundary list (it never names LangChain and names five neighbours), the SWE-agent Lite figure's framing, the METR 26pp/8pp figures and the elicitation attribution, the lm-evaluation-harness quote source and Zenodo DOI, the MMLU table label and the HELM read-out description, the Open LLM Leaderboard quote, the tau-bench authorship, the test-harness definitions' sources, the TechTarget URL and the harness.io founding wording; added OpenRouter's own dated app rankings with its stated caveats, identified omp as Oh My Pi, Command Code and Freebuff, and added Harness-Bench's measured 23.8-point harness gap, HELM's maintenance mode and simple-evals' deprecation.
Version 6 · Oct 1, 2026, 10:58 AM