Revision History
Benchmark (AI) · 11 revisions
Sizes are character counts of the article source. The signed number is the change from the previous revision.
Recent edit summaries
Detailed summaries recorded by editors. Generic maintenance summaries are omitted here; every recorded revision remains below. Dates describe the edit, not necessarily the event it covers.
- Correction: the claim that LMArena disputed the paper's figures now cites LMArena's own response post rather than the paper it responds to
Version 11 · Aug 4, 2026, 11:11 PM
- Add 'Benchmarking as an industry' section: commercialization of AI evaluation (LMArena/Arena, Scale SEAL, Artificial Analysis, Mercor, Epoch AI), independence controversies; extend References 37-46 and See also
Version 10 · Aug 4, 2026, 10:34 PM
- Correction: replace unsupported and volatile benchmark claims with source-bounded guidance; remove contradictory BetterBench counts, qualify LLM-judge self-enhancement evidence, correct GAIA and Tau-bench publication records, and add construct-validity, uncertainty, contamination, resilience, and agent-evaluation coverage.
Version 9 · Jul 29, 2026, 05:00 AM
- Added 6 contextual internal links
Version 8 · Jul 23, 2026, 03:00 PM
- Add internal link to /wiki/lm_evaluation_harness (harness cluster)
Version 6 · Jul 14, 2026, 01:47 AM