AA-Briefcase
AA-Briefcase is a private benchmark from Artificial Analysis that tests AI agents on long-horizon office work. Models work through multi-week business projects built from thousands of input files and must hand in finished deliverables such as spreadsheet models, slide decks, memos and PDF reports. Each deliverable is graded three ways: binary rubric checks for correctness, a pairwise comparison of analytical quality, and a pairwise comparison of presentation quality. The three results are combined into a single AA-Briefcase Elo.[1][3] Artificial Analysis announced the benchmark on 18 June 2026. On 4 September 2026 it became the most heavily weighted agentic evaluation in the Artificial Analysis Intelligence Index v4.2, at 15%.[1][4] The current version, v1.1, changed only how the Elo ratings are fitted.[3] As of 29 September 2026, Claude Opus 5.5 at max effort led the leaderboard at 1822 Elo, with Claude Sonnet 5.5 at max effort second at 1811.[2] Anthropic, SpaceXAI and StepFun all put AA-Briefcase scores in their model launch materials.[17][18][19][20][21][22]
Overview
| Property | Detail |
|---|---|
| Developer | Artificial Analysis[1] |
| Announced | 18 June 2026[1] |
| Current version | v1.1 (Elo ratings fitted with a Crowd-BT model)[3] |
| Scored content | 4 private scenarios, 91 tasks in total[1][3] |
| Public example | AA-Briefcase-Lite: one week of a fifth scenario, released on Hugging Face under Apache 2.0 and not scored[6] |
| Deliverables | Excel, PowerPoint, PDF, Word and other formats such as HTML and LaTeX[2] |
| Harness | Stirrup, Artificial Analysis' open-source agent framework, in a week-scoped, offline E2B sandbox[3] |
| Turn limit | 500 turns per task[3] |
| Repeats | 1 run per task[3] |
| Grading | Binary rubric checks plus pairwise analytical-quality and presentation judgements, decided by three-model judge panels[3] |
| Headline metric | AA-Briefcase Elo, anchored so that GPT-5.5 (medium) = 1000[3] |
| Intelligence Index weight | 15% since v4.2 (4 September 2026), in the 30% Agents category[3][4] |
Background
Artificial Analysis built AA-Briefcase because models are increasingly used for "complex long-horizon knowledge work tasks". The company says the benchmark tries to mirror that use: tasks build week by week, share institutional context, and end in realistic company deliverables, which it contrasts with "single, disconnected prompts".[1] Artificial Analysis says the tasks were "developed over months by experts across data science, product management and corporate strategy from companies including Google, McKinsey & Company and Boston Consulting Group".[1]
Artificial Analysis also runs GDPval-AA, its implementation of OpenAI's GDPval dataset of tasks across 44 occupations, whose test set is not private.[2][3] AA-Briefcase differs in keeping its task set private and in grouping tasks into continuous projects that share a pool of source material. The index update that brought AA-Briefcase in, v4.2, was billed as having "more complex and realistic tasks, and more private test sets to prevent gaming".[4]
Design
Scenarios
Each scenario is a multi-week project. The agent works through it in sequence, and each week has two to five tasks.[3] The four scored scenarios are held out. A fifth, public scenario shows the format and does not count toward any score.[1]
| Scenario | Status | Description (Artificial Analysis) | Example tasks |
|---|---|---|---|
| Data Science | Held out | Quantitative work on imperfect datasets, turned into recommendations for a business audience | Transaction data cleaning and reconciliation; forecast modelling and feature engineering; data engineering and schema design |
| Product Management | Held out | The work of a product leader in a complex business | Competitive teardowns and market research; MVP feature prioritization; PRD writing; go-to-market and pilot planning |
| Banking Operations | Held out | A retail-banking branch-network transformation | Branch traffic analysis by channel; staff time-allocation survey design; branch servicing financial model; mortgage journey mapping |
| Heavy Industry Strategy | Held out | Strategic and investment decisions in an asset-heavy industrial business | 10-year commodity supply-demand and price model; asset-level operating model; policy impact matrix on retain-vs-divest economics; precedent M&A comps |
| Due Diligence | Public (AA-Briefcase-Lite) | Outside-in due diligence on an agricultural food producer for a private equity firm | Market structure map; market sizing and cage-ban transition model; target assessment deck; briefing video |
Source: Artificial Analysis launch article.[1]
Source material
Each task gives the agent hundreds of input files. Across the benchmark there are nearly 2,000 source files, and the email and Slack exports alone hold more than 3,500 emails and 25,000 Slack messages.[1] The methodology page lists the file types as Slack exports, spreadsheets, PDFs, interview transcripts, market research, standards documents, app-store pages, board materials, emails and other business records. The mix includes real, augmented and synthetic material.[3] Artificial Analysis describes the sources as "fragmented, messy, and often contain realistic contradiction". In the public AA-Briefcase-Lite scenario, "Critical Insight" rubric checks test whether a deliverable reconciles contradictions deliberately planted across sources.[6]
Each scenario's pool has shared files, available for the whole project, and week-specific files. Later-week tasks may also get "standardized base-case files", the same reference work products given to every model. These let each task run on its own while keeping the project continuous.[3]
Deliverables
On the leaderboard, results are broken down by the file type of the required deliverable: Excel, PowerPoint, PDF, Word, and "Other" for formats such as HTML and LaTeX.[2] For a model that completed the full set, the page's per-file-type data counts 38 Excel tasks, 22 PowerPoint, 17 Word, 12 PDF and 2 other, which adds up to the 91 scored tasks.[2]
How models are run
Models run through Stirrup, Artificial Analysis' open-source agent framework, inside a week-scoped E2B sandbox.[3] The sandbox has no internet access, so the agent can only use the files provided. It comes with Python 3.13 and a data and document stack (pandas, python-docx, python-pptx, openpyxl, PyMuPDF and others), plus tools such as LibreOffice, Pandoc, Tesseract, FFmpeg and TeX Live. Individual commands are stopped after 20 minutes.[3]
The agent gets one code-execution tool, a view-image tool if the model accepts images, and two finishing tools. The finish tool submits a summary and the absolute paths of every deliverable. The abandon_task_finish tool lets the agent give up with a reason. The methodology page says the agent calls it "only when it concludes the task is genuinely impossible", and the system prompt adds: "Do not use it to escape difficulty."[3] Each agent gets up to 500 turns per task, and the agent cannot talk to a user during the run.[3] According to the dataset card, Stirrup summarizes earlier history when the context window fills, so long tasks can continue.[6]
Tasks within a scenario share files and context across weeks, but models currently do each task in a separate run and do not see their own earlier submissions.[1][3] Each task is run once.[3]
Grading
Three kinds of check
Every task is graded in three ways.[1]
| Check | Type | Question asked |
|---|---|---|
| Rubric | Binary pass or fail per check | Did the model follow the task instructions, find requirements hidden across source files, use the correct evidence and reach the right conclusions? |
| Analytical quality | Pairwise comparison | Compared with another model's submission, which deliverable is more thorough, analytically rigorous and well supported? |
| Presentation | Pairwise comparison | Compared with another model's submission, which one is more professionally presented? |
The published AA-Briefcase-Lite week splits rubric checks into two types. "Accuracy" checks cover single-source facts and instruction following. "Critical Insight" checks require reconciling contradictory sources or catching a deliberate trap. Rubric checks get no partial credit. Pairwise checks return a preferred submission or a tie.[6] The rubric judge's prompt tells it to use "only evidence from the submitted artifact content". It sees the task, the rubric item and the submission, never the source files, so it judges citations on whether they are "present, specific, and well-formed" rather than on whether it can check them against the sources.[3][6]
Judge panels
Rubric grading and the two pairwise comparisons each use a panel of three judge models. Artificial Analysis says a panel reduces "bias toward submissions from the same model or model family". Each rubric verdict or pairwise match is decided by one judge sampled from the panel, with sampling balanced across checks and matches. A given rubric check is always graded by the same judge.[3]
| Panel | Judges (current methodology) |
|---|---|
| Rubric grading | Claude Opus 4.8 (max effort), GPT-5.5 (high reasoning), Gemini 3.1 Pro Preview (high reasoning) |
| Pairwise (analytical quality and presentation) | Claude Opus 5 (high effort), GPT-5.6 Sol (medium reasoning), Gemini 3.8 Flash (high reasoning) |
Source: Artificial Analysis methodology page, accessed 29 September 2026.[3]
The pairwise panel is newer than the rubric panel. Intelligence Index v4.3.1, in September 2026, "refreshed the pairwise judge panels to current model versions", moving the AA-Briefcase and GDPval-AA pairwise comparisons to Claude Opus 5, GPT-5.6 Sol and Gemini 3.8 Flash. The rubric panels were not changed.[3]
From judgements to Elo
The headline AA-Briefcase Elo combines analytical-quality Elo, presentation Elo and rubric pass rate. Rubric performance is converted into Elo through "synthetic head-to-head matches", and everything is combined by maximum-likelihood Elo aggregation.[3] The scale is anchored so that GPT-5.5 (medium) sits at 1000, and Elo scores and confidence-interval bounds are clamped at 0.[2][3] Leaderboard captures from the day after launch onward show GPT-5.5 (medium) fixed at 1000 with a zero-width interval.[23]
In v1.1 the ratings in each scope are fitted with a Crowd-BT model.[3] Crowd-BT, from a 2013 paper by Xi Chen, Paul N. Bennett, Kevyn Collins-Thompson and Eric Horvitz, extends the Bradley-Terry model of pairwise preferences by explicitly modelling how reliable each annotator is.[7] Artificial Analysis fits the annotator-quality term per scope on past AA-Briefcase judgements. For the rubric scope it notes that grading is "decided by a deterministic comparison rather than a judge" and treats that scope accordingly.[3]
Versions
| Version or change | When | What changed |
|---|---|---|
| v1 | 18 June 2026 launch | Original release; 91 tasks in 4 private scenarios[1]; GPT-5.5 (medium) shown at 1000 from the first archived capture[23] |
| Index v4.2 update | 4 September 2026 | Artificial Analysis "improved our sampling and re-anchored the Elo scale" for AA-Briefcase and GDPval-AA v2, "making ratings more stable as new models are added"; AA-Briefcase joined the Intelligence Index[4] |
| Index v4.3.1 | September 2026 | Pairwise judge panel refreshed to Claude Opus 5, GPT-5.6 Sol and Gemini 3.8 Flash; rubric panel unchanged[3] |
| v1.1 (Index v4.3.2) | September 2026 | "v1.1 changes only how Elo ratings are fitted": each pairwise scope is fitted with Crowd-BT; anchor unchanged at GPT-5.5 (medium) = 1000[3] |
The methodology page dates v4.3.1 and v4.3.2 only to "September 2026". Web Archive captures of the leaderboard show it still titled "AA-Briefcase" on 17 September and retitled "AA-Briefcase v1.1" by 21 September 2026.[23] Artificial Analysis says that under v1.1 "Elo scores shift, rank ordering is largely preserved".[3]
Because of these changes, AA-Briefcase figures from different dates are not directly comparable, even for the same model configuration.[3][23] The table below tracks two configurations through Artificial Analysis' articles and archived captures of its leaderboard.
| Date | Source | Claude Fable 5 (max) | Claude Opus 5 (max) |
|---|---|---|---|
| 18 June 2026 | Launch data[1] | 1587 | not yet released |
| 24 July 2026 | Opus 5 launch article[9] | 1574 | 1720 |
| 18 August 2026 | Leaderboard capture[23] | 1574 | 1714 |
| 1 September 2026 | Fable 5.1 article[10] | not stated | 1685 |
| 7 September 2026 | Leaderboard capture, after the v4.2 re-anchoring[23] | 1534 | 1647 |
| 21 September 2026 | Leaderboard capture, first seen as v1.1[23] | 1543 | 1673 |
| 29 September 2026 | Live leaderboard (v1.1)[2] | 1541 | 1673 |
Role in the Intelligence Index
AA-Briefcase joined the Artificial Analysis Intelligence Index in v4.2, announced on 4 September 2026, at 15% weight in the Agents category.[4] In v4.3, the other Agents evaluations are GDPval-AA v2 (10%) and AutomationBench-AA (5%).[3][5] AA-Omniscience also carries 15%, but split into accuracy (10%) and non-hallucination (5%), so no single evaluation outweighs AA-Briefcase.[3] Artificial Analysis said v4.2 raised the private, held-out share of the index to 40%, "double the figure from v4.1", and named AA-Briefcase first among the held-out sets.[4]
For the index, a model's combined v1.1 Elo is "frozen at the time of a model's addition" and normalized as clamp((Elo - 500) / 2000). GDPval-AA v2.1 uses the same mapping. By this formula an Elo of 500 counts as zero and 2500 as full marks, so Opus 5.5's 1822 maps to about 0.66. Artificial Analysis says it "may update the reference parameters as models progress against the evaluation".[3]
AA-Briefcase also feeds Artificial Analysis' occupation-weighted Capability Indices. Version 1.1 of those indices, announced on 14 September 2026, added it next to GDPval-AA v2 in the "Agentic Knowledge Work" component of all six indices: Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering and Economics. That component carries 20% to 35% of each index.[15]
AA-Briefcase-Lite
AA-Briefcase-Lite, the public example, was published on Hugging Face on 18 June 2026 under the Apache 2.0 license.[6] It holds one week of a fifth scenario, which the dataset card describes as "a smaller, less challenging scenario than the full benchmark". It is not part of the scored leaderboard.[6] The scenario is a commercial due-diligence engagement. The agent assists a vice president at Halberd Capital Partners, a fictional US mid-market private-equity firm, in evaluating Aurora Eggs Ltd, a fictional New Zealand egg producer. The companies are invented but the scenario is "grounded in real data from the New Zealand egg market".[6]
| Task | Workstream | Deliverable |
|---|---|---|
| w1_t1 | Market structure and competitive landscape | One-page market_overview.tex plus PDF |
| w1_t2 | Market sizing, forecast and cage-ban transition model | market_model.xlsx |
| w1_t3 | Target assessment, opportunities and risks | target_assessment.pptx |
| w1_t4 | Preliminary findings briefing video with subtitles | briefing.mp4 plus briefing.srt |
Source: AA-Briefcase-Lite dataset card.[6]
The source pool has 67 sources, organized into shared and week-specific folders. The release includes 63 grading checks, the grader's rubric, source-to-check mappings and evidence chains, the agent and judge prompts verbatim, and example submissions from six models: Claude Opus 4 (reasoning), Claude Fable 5, GLM 5.2 (max), Gemini 3.1 Pro, GPT-5.5 (xhigh) and o3.[6] The leaderboard page notes that 32 of the 67 sources are relevant to the example task shown.[2]
Results
At launch (June 2026)
At launch, Claude Fable 5 led, followed by Claude Opus 4.8 (max) and GLM-5.2 (max). Artificial Analysis called GLM-5.2 "the clear leader among open-weight models".[1]
| Model (configuration) | Elo | Analytical quality Elo | Presentation Elo | Rubric pass rate | Cost per task |
|---|---|---|---|---|---|
| Claude Fable 5 (max, Opus 4.8 fallback) | 1587.2 | 1774 | 1493 | 56.0% | $31.18 |
| Claude Opus 4.8 (max) | 1356.0 | 1367 | 1493 | 38.7% | $10.40 |
| GLM-5.2 (max) | 1265.5 | 1346 | 1292 | 36.1% | $2.40 |
| GPT-5.5 (xhigh) | 1158.7 | 1227 | 1123 | 33.4% | $3.68 |
| MiniMax-M3 | 1116.2 | 1090 | 1200 | 30.3% | $1.24 |
| Claude Sonnet 4.6 (max) | 1081.4 | 1012 | 1202 | 28.9% | $2.41 |
| DeepSeek V4 Pro (max) | 935.8 | 902 | 950 | 23.9% | $0.10 |
| Gemini 3.5 Flash (high) | 870.2 | 703 | 891 | 27.7% | $4.91 |
| DeepSeek V4 Flash (max) | 836.0 | 757 | 885 | 18.7% | $0.04 |
| Gemini 3.1 Pro Preview | 445.2 | 267 | 336 | 12.4% | $0.82 |
Source: chart data in the launch article, "data as at 18 June 2026". Cost per task is the total cost to run the benchmark divided by 91 tasks, from token usage and list prices with representative cache-hit rates.[1] The table is a selection of the 19 models charted.
Findings Artificial Analysis reported at launch:[1]
- Far from saturated. Fable 5 led on rubric pass rate but met every rubric criterion on only 3% of tasks. On 31 of the 91 tasks, no model scored above 50%.
- Failure modes change with capability. Weaker models more often failed at execution: missing input files, submitting unusable deliverables, or submitting nothing. Stronger models more often missed requirements, including requirements hidden across source files. Incorrect or unfinished analysis and formatting errors were common at every tier.
- More required files, lower pass rates. For models averaging at least 30% on the rubric, pass rates fell from about 55% on checks that need only the prompt to about 40% on checks that need five or more files.
- Looking at the output helps presentation. Fable 5 and Opus 4.8 (max), the two leaders on presentation Elo, averaged 21 and 12 view-image calls per task. GPT-5.4 Mini averaged 2 and Gemini 3.1 Pro Preview about 0.1, "often submitting files they never visually reviewed".
- Time and tokens. Fable 5 averaged 139k output tokens per task. Gemini 3.5 Flash used the most, 146k, while scoring about 720 Elo lower. Opus 4.8 (max) averaged about 24 minutes per task and GLM-5.2 (max) about 19. MiniMax-M3 took about 26 minutes and still finished 240 Elo behind Opus 4.8. Artificial Analysis found no strong correlation between turn count and score.
- Some models over- or under-perform their general scores. MiniMax-M3 and GLM-5.2 did better on AA-Briefcase than their Intelligence Index scores would suggest. Gemini 3.5 Flash and Gemini 3.1 Pro Preview did worse.
Measured cost per task spanned more than 800 times, from about $0.04 for DeepSeek V4 Flash (max) to more than $31 for Fable 5. By Artificial Analysis' account, GLM-5.2 (max) came within about 90 Elo of Opus 4.8 (max) for under a quarter of the cost.[1]
July to mid-September 2026
- GPT-5.6 Sol (July 2026). Artificial Analysis reported that GPT-5.6 Sol (max) ranked second behind Fable 5 and had "the highest recorded Presentation Elo" of any model. Fable 5 still led on rubric score (56% against 42%) and analytical-quality Elo (1764 against 1592).[8]
- Claude Opus 5 (24 July 2026). Opus 5 at max effort scored 1720, 146 points ahead of Fable 5 (1574), at $17.79 per task against $22.30 for Fable 5. Its analytical-quality Elo of 2016 was nearly 300 above Fable 5. Its presentation Elo of 1628 remained about 40 behind GPT-5.6 Sol (max, 1666). The top three effort settings each averaged more than 25 minutes per task.[9]
- Claude Fable 5.1 (1 September 2026). Fable 5.1 scored 1694 against Opus 5's 1685, which Artificial Analysis called "effectively tied". Fable 5.1 was ahead on analytical quality (2025 against 1980) and behind on presentation (1495 against 1572).[10]
- GPT-6 Astra (9 September 2026). Artificial Analysis measured a gain of about 90 Elo over GPT-5.6 Sol, driven by rubric score and analytical quality. Presentation Elo fell, and GPT-5.6 Sol (max) still led all models there.[11]
- Grok 4.7 (21 September 2026). Grok 4.7 (xhigh) scored 1657, up 111 from Grok 4.6 (high). Its analytical-quality Elo was 1994 and its presentation Elo 1499.[12]
- GPT-6 Sol and Luna (22 September 2026). GPT-6 Sol was level with its predecessor on AA-Briefcase v1.1, and GPT-6 Luna dropped about 45 Elo. Artificial Analysis said its team had "manually inspected hundreds of model outputs" and that the knowledge-work regressions "tend to be driven by reduced presentation quality and deliverables that omit rubric elements".[14]
Current leaderboard (v1.1)
Claude Opus 5.5 took the lead at its 22 September 2026 launch. Artificial Analysis said it led on analytical quality and presentation and sat "just behind Fable 5.1 for rubric-based scoring". The firm said it was the first time an Anthropic model had surpassed GPT-5.6 Sol on presentation quality.[13] Six days later, Claude Sonnet 5.5 at max effort scored 1811, which Artificial Analysis described as parity with Opus 5.5, "albeit with significantly higher token usage".[16]
| Rank | Model (configuration) | Elo (95% CI) | Analytical quality Elo | Presentation Elo | Rubric pass rate |
|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 (max) | 1822 (±12) | 2207 | 1710 | 61.2% |
| 2 | Claude Sonnet 5.5 (max) | 1811 (±11) | 2161 | 1694 | 62.4% |
| 3 | Claude Opus 5.5 (xhigh) | 1780 (±11) | 2151 | 1640 | 60.6% |
| 4 | Claude Sonnet 5.5 (xhigh) | 1746 (±10) | 2051 | 1597 | 63.2% |
| 6 | Claude Fable 5.1 (max) | 1678 (±10) | 1999 | 1486 | 61.5% |
| 7 | Claude Opus 5 (max) | 1673 (-10/+11) | 1965 | 1560 | 57.2% |
| 9 | Grok 4.7 (xhigh) | 1657 (±9) | 1994 | 1499 | 55.4% |
| 14 | Qwen3.8 Max (0902) | 1626 (±11) | 1989 | 1394 | 58.0% |
| 17 | Muse Spark 1.3 (max) | 1587 (-10/+11) | 1721 | 1523 | 58.7% |
| 19 | GPT-6 Astra (max) | 1569 (-10/+11) | 1777 | 1535 | 52.0% |
| 26 | GLM-5.3 (max) | 1512 (±9) | 1787 | 1349 | 51.3% |
| 28 | Kimi K3 (max) | 1505 (-7/+8) | 1680 | 1446 | 51.0% |
| 32 | GPT-5.6 Sol (max) | 1487 (-8/+9) | 1559 | 1653 | 41.8% |
| 34 | GPT-6 Sol (max) | 1483 (±10) | 1652 | 1473 | 46.9% |
| 41 | Step 5 Preview | 1432 (±14) | 1574 | 1388 | 47.8% |
| 49 | Claude Sonnet 5 (max) | 1359 (±7) | 1419 | 1412 | 42.3% |
| 102 | GPT-5.5 (medium), anchor | 1000 | 1000 | 1000 | 26.6% |
Source: Artificial Analysis AA-Briefcase v1.1 leaderboard, accessed 29 September 2026. The leaderboard lists 202 model configurations. Sub-scores come from the leaderboard's per-configuration score breakdown.[2]
Sub-scores show different strengths. GPT-5.6 Sol (max) has the third-highest presentation Elo among the charted configurations, behind only Opus 5.5 (max) and Sonnet 5.5 (max), but ranks 32nd overall because of its lower rubric pass rate and analytical-quality Elo. Grok 4.7 and Qwen3.8 Max (0902) are close to Fable 5.1 and Opus 5 on analytical quality (1994 and 1989, against 1999 and 1965) but further behind on presentation (1499 and 1394).[2]
Use in model launches
Within weeks of launch, AA-Briefcase scores were appearing in other companies' model launch tables.
| Launch (date) | Publisher | Label used | Scores shown |
|---|---|---|---|
| Claude Opus 5 (24 Jul 2026) | Anthropic (system card, section 8.13.5) | "AA-Briefcase" | Opus 5 1720; Opus 4.8 1346; Fable 5 1574; GPT-5.6 Sol 1505[24] |
| Grok 4.6 (12 Aug 2026) | SpaceXAI | "AA-Briefcase" | Grok 4.6 High 1577; Grok 4.5 High 1313; GPT-5.6 Sol Max 1502; Fable 5 Max 1574[20] |
| Claude Fable 5.1 (1 Sep 2026) | Anthropic (system card, section 8.15.4) | "AA-Briefcase" | Fable 5.1 1694; Fable 5 1572; Opus 5 1685; GPT-5.6 Sol 1502[25] |
| Step 5 Preview (Sep 2026) | StepFun | "AA-Briefcase v1.1" | Step 5 Preview (High) 1433; GLM-5.3 (Max) 1526; Kimi K3 (Max) 1511; GPT-6 Astra (Max) 1569; Claude Fable 5.1 (Max) 1678; Claude Opus 5 (Max) 1673[22] |
| Grok 4.7 (21 Sep 2026) | SpaceXAI | "Multi-hour office work: AA Briefcase v1.1" | Grok 4.7 1,657; Grok 4.6 1,546; GPT-5.6 Sol 1,487; Fable 5.1 1,678[21] |
| Claude Opus 5.5 (22 Sep 2026) | Anthropic (system card, section 8.14.4) | "AA-Briefcase v1.1" | Opus 5.5 1822; Opus 5 1673; Fable 5.1 1678; GPT-6 Astra 1569[19] |
| Claude Sonnet 5.5 (28 Sep 2026) | Anthropic (launch page and system card) | "AA-Briefcase v1.1" | Sonnet 5.5 1811; Sonnet 5 1359; Opus 5.5 1822; GPT-6 Sol 1483[17][18] |
SpaceXAI's Grok 4.6 table says third-party scores are "the best of self-reported or publicly available results".[20] Anthropic's system cards say the AA-Briefcase evaluations "were run independently by Artificial Analysis".[18][19][24] Values printed at launch can differ from the live leaderboard by anything from one point to more than ten: StepFun printed 1526 for GLM-5.3 (max), which the live leaderboard shows at 1512.[22][2] SpaceXAI printed 1577 for Grok 4.6 at launch, and a capture of the leaderboard six days later showed Grok 4.6 (high) at 1578.[20][23]
Claude Sonnet 5.5
Anthropic's Sonnet 5.5 launch table reports Sonnet 5.5 at 1811 on AA-Briefcase v1.1, against 1359 for Sonnet 5, 1822 for Opus 5.5 and 1483 for GPT-6 Sol. Two footnotes are attached.[17]
- Footnote 3 says Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment of Sonnet 5.5 on the Claude Platform, "which we found to have a bug that could degrade responses to requests that use structured outputs". Anthropic expected any effect "to be small and to understate its performance", and said the bug has been fixed.[17]
- Footnote 4, on the GPT-6 Sol figure, says OpenAI had recently fixed a bug that degraded image understanding in GPT-6 Sol, and that official AA-Briefcase v1.1 and GDPval-AA v2.1 scores "may not have been updated yet". It adds that "Artificial Analysis does not expect major impacts".[17]
Anthropic also says that on AA-Briefcase, Sonnet 5.5 at medium effort "bests Sonnet 5's best score for about one ninth of the cost per task".[17] The live leaderboard is consistent with the scores in that claim: Sonnet 5.5 (medium) is at 1461 and Sonnet 5 (max) at 1359.[2] The Sonnet 5.5 system card adds 1746 for xhigh effort and says that setting uses about 61% fewer output tokens than max.[18] Artificial Analysis' own launch article used the same 1811 figure, mentioned the structured-outputs bug, and said it "will be re-running relevant evaluations soon".[16]
Caveats
- Private tasks. The scored scenarios, rubrics and input files are private to limit contamination and gaming. Outsiders therefore cannot reproduce official scores. The public Lite scenario shows the format but is described as easier than the scored set.[1][6]
- Independent task runs. Models do not carry their own earlier work from week to week, so the benchmark does not yet test whether an agent can build on its own previous deliverables.[1][3]
- Model judges. Every rubric verdict and pairwise judgement comes from an LLM judge. The judges come from the same model families being ranked, which is why Artificial Analysis uses three-judge panels.[3]
- Single run per task. Each configuration is run once per task. The published 95% confidence intervals for the leading configurations are roughly ±7 to ±14 Elo.[2][3]
- Moving scale. Re-anchoring in v4.2 and the Crowd-BT refit in v1.1 changed published Elo values for models that were not re-run, so launch-day figures from different months should not be compared directly.[3][4][23]
See also
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20Artificial Analysis. "Announcing AA-Briefcase: a frontier knowledge work evaluation." June 18, 2026. artificialanalysis.ai/...aa-briefcase (launch figures read from the page's embedded chart data, "Data as at 18 June 2026").
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13Artificial Analysis. "AA-Briefcase v1.1: Agentic Knowledge Work Benchmark" (leaderboard and charts). artificialanalysis.ai/...aa-briefcase. Accessed September 29, 2026.
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33 ^34 ^35 ^36 ^37 ^38 ^39Artificial Analysis. "Intelligence Benchmarking" methodology page (Intelligence Index v4.3.2; AA-Briefcase v1.1 section and version history). artificialanalysis.ai/...intelligence-benchmarking. Accessed September 29, 2026.
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Artificial Analysis. "Announcing Artificial Analysis Intelligence Index v4.2." September 4, 2026. artificialanalysis.ai/...s-intelligence-index-v4-2
- ^Artificial Analysis. "Announcing the Artificial Analysis Intelligence Index v4.3." September 7, 2026. artificialanalysis.ai/...s-intelligence-index-v4-3
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11Artificial Analysis. "AA-Briefcase-Lite" dataset card. Hugging Face. huggingface.co/...AA-Briefcase-Lite. Accessed September 29, 2026.
- ^Chen, Xi; Bennett, Paul N.; Collins-Thompson, Kevyn; Horvitz, Eric. "Pairwise Ranking Aggregation in a Crowdsourced Setting." WSDM '13, Rome. erichorvitz.com/crowd_pairwise.pdf (read via Internet Archive capture of August 30, 2024).
- ^Artificial Analysis. "GPT-5.6 benchmarks across Intelligence, Speed and Cost." July 9, 2026. artificialanalysis.ai/...gpt-5-6-has-landed
- ^1 ^2Artificial Analysis. "Claude Opus 5: the new leader in agentic knowledge work." July 24, 2026. artificialanalysis.ai/...er-agentic-knowledge-work
- ^1 ^2Artificial Analysis. "Claude Fable 5.1 tops the Artificial Analysis Intelligence Index." September 1, 2026. artificialanalysis.ai/...claude-fable-5-1
- ^Artificial Analysis. "Benchmarking GPT-6 Astra." September 9, 2026. artificialanalysis.ai/...benchmarking-gpt-6-astra
- ^Artificial Analysis. "Benchmarking Grok 4.7." September 21, 2026. artificialanalysis.ai/...benchmarking-grok-4-7
- ^Artificial Analysis. "Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index." September 22, 2026. artificialanalysis.ai/...claude-opus-5-5
- ^Artificial Analysis. "GPT-6 Sol and Luna push the cost efficiency frontier." September 22, 2026. artificialanalysis.ai/...-cost-efficiency-frontier
- ^Artificial Analysis. "Announcing Artificial Analysis Capability Indices v1.1." September 14, 2026. artificialanalysis.ai/...s-capability-indices-v1-1
- ^1 ^2Artificial Analysis. "Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index." September 28, 2026. artificialanalysis.ai/...claude-sonnet-5-5
- ^1 ^2 ^3 ^4 ^5 ^6Anthropic. "Introducing Claude Sonnet 5.5." September 28, 2026. anthropic.com/claude-sonnet-5-5
- ^1 ^2 ^3 ^4Anthropic. "Claude Sonnet 5.5 System Card" (Table 8.1.A; section 8.14.4, AA-Briefcase). September 2026. anthropic.com/claude-sonnet-5-5-system-card
- ^1 ^2 ^3Anthropic. "Claude Opus 5.5 System Card" (Table 8.1.A; section 8.14.4, AA-Briefcase). September 2026. anthropic.com/claude-opus-5-5-system-card
- ^1 ^2 ^3 ^4SpaceXAI. "Introducing Grok 4.6." August 12, 2026. x.ai/...grok-4-6 (read via archived copy: web.archive.org/...grok-4-6)
- ^1 ^2SpaceXAI. "Introducing Grok 4.7." September 21, 2026. x.ai/...grok-4-7 (read via archived copy: web.archive.org/...grok-4-7)
- ^1 ^2 ^3StepFun. "Step 5 Preview: Advancing the Pareto Frontier." September 2026. stepfun.com/step-5-preview. Accessed September 29, 2026.
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Internet Archive captures of artificialanalysis.ai/...aa-briefcase: June 19, 2026 (web.archive.org/...aa-briefcase); August 18, 2026 (web.archive.org/...aa-briefcase); September 7, 2026 (web.archive.org/...aa-briefcase); September 17, 2026 (web.archive.org/...aa-briefcase); September 21, 2026 (web.archive.org/...aa-briefcase).
- ^1 ^2Anthropic. "System Card: Claude Opus 5" (Table 8.1.A; section 8.13.5, AA-Briefcase). July 24, 2026. www-cdn.anthropic.com/...s%205%20System%20Card.pdf
- ^Anthropic. "System Card: Claude Fable 5.1 & Claude Mythos 5.1" (Table 8.1.A; section 8.15.4, AA-Briefcase). September 1, 2026. www-cdn.anthropic.com/...205.1%20System%20Card.pdf
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,916 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: xg12 V4 independent verification 29 Sep 2026: 52 sources, ~230 claims (both clusters); 1 material (misattributed comparability caveat) + 9 minor fixed.
Cite this page: AI Wiki. "AA-Briefcase." aiwiki.ai, updated 29 Sept 2026, fact-checked 29 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/aa_briefcase