AutomationBench
AutomationBench is a benchmark for AI agents, built by the automation company Zapier, that tests whether a model can finish a multi-step business workflow across several simulated software applications. An agent gets one plain-language request (for example, close a deal in the CRM and email the right team), has to find the relevant REST API endpoints on its own, follow company policies hidden in spreadsheets or inboxes, and leave the simulated apps in the correct final state. Grading is done by deterministic checks on that end state, with no LLM-as-a-judge.[1][2][3] Zapier announced the benchmark on April 20, 2026, and its authors, Daniel Shepard and Robin Salimans of Zapier, posted the accompanying paper to arXiv on April 21, 2026 (UTC).[1][2] At launch the best model scored 9.9%. By late September 2026 the top scores on Zapier's leaderboard were in the low-to-mid 40s, after Anthropic, OpenAI, Google and DeepSeek had started quoting the benchmark in model launch materials.[1][3][13][15][16][17][18] Artificial Analysis runs its own variant, AutomationBench-AA, which scores partial completion and since September 2026 has been part of its Intelligence Index.[9][10][11]
Overview
| Property | Value |
|---|---|
| Developer | Zapier (authors Daniel Shepard and Robin Salimans)[1] |
| Announced | April 20, 2026 (Zapier blog); arXiv:2604.18934 submitted April 21, 2026[1][2] |
| Current version | 1.0.6 (changelog dated July 31, 2026)[5] |
| Domains | Sales, Marketing, Operations, Support, Finance, HR[1] |
| Simulated apps and endpoints | 47 apps, about 500 API endpoints[1][3] |
| Public task set | 600 tasks (100 per domain), plus a 200-task "simple" baseline domain that is not scored[4] |
| Private task set | Held out; described as "600+" tasks by Zapier; 657 tasks in the release 1.0.6 split used by Artificial Analysis[3][11] |
| Agent tools | Two tools: search (BM25 over API schemas) and execute (a simulated HTTP request)[1][3] |
| Step limit | 50 steps per task[1][4] |
| Headline metric | task_completed_correctly: a task passes only if every assertion holds[3][4] |
| Grader | Programmatic assertions on simulated app state; no LLM judge[1][2] |
| Code license | MIT (original Zapier code); third-party API schema representations are not claimed as original works[6] |
| Leaderboard | zapier.com/benchmarks[3] |
History
The GitHub repository zapier/AutomationBench was created on February 23, 2026, and its first commit, at version 0.1.0, followed on March 5, 2026. An April 2 commit, "Rename benchmark to automationbench and update to latest benchmark version," set the current name and moved the package to version 0.1.1. The finance and HR domains were added between April 16 and April 18, 2026.[7]
Zapier announced the benchmark in a blog post by Anna Marie Clifton, who leads Zapier Agents and AI at the company, on April 20, 2026. The post calls it "the first open benchmark that scores AI models on end-to-end business workflow execution", a claim that is Zapier's own.[2] According to the post, Zapier "initially built AutomationBench for internal use, to help evaluate which models to deploy across Zapier" and released it once it "proved useful enough." It launched with a public task set, and Zapier said model providers could request verified evaluation on the private set.[2] The paper, titled simply "AutomationBench," was submitted to arXiv on April 21, 2026 (UTC).[1] The same day, Prime Intellect announced that it was hosting the benchmark on its Environments Hub, writing that "across 6 domains, 47 tools, and 600 tasks, frontier models all score under 10%."[8] The benchmark is built on Prime Intellect's Verifiers framework, and the authors say they ran reinforcement learning with verifiable rewards experiments in Prime Intellect's Lab to find reward-hacking strategies and harden the reward function.[1] The paper's acknowledgements thank Prime Intellect, Anthropic and OpenAI for early feedback, and Mike Knoop for feedback and benchmark expertise from his work on the ARC Prize.[1]
Artificial Analysis launched AutomationBench-AA on July 6, 2026, in partnership with Zapier, running on the private task subset.[9] On September 7, 2026, it added AutomationBench-AA to version 4.3 of its Intelligence Index at a 5% weight, replacing τ³-Banking.[10][11]
Design
Domains and environments
The paper says the six domains "are the most popular types of workflows Zapier customers use to automate their businesses," and that a pool of about 500 API endpoints across 47 apps is shared across them.[1] Zapier's leaderboard page describes this as "47 real tools across six business functions."[3] The README describes the coverage as 47 simulated SaaS tools, with domains ranging from CRM and lead management (Sales) to ad performance and brand monitoring (Marketing), vendor workflows and compliance (Operations), ticket routing and SLA monitoring (Support), accounts payable and receivable and bookkeeping (Finance), and recruitment, onboarding, time off and payroll (HR).[4] Artificial Analysis names Gmail, Google Sheets, Slack, Salesforce, Zendesk, Jira and HubSpot among the simulated apps. Its July launch article referred to "40 simulated app environments," a different count from Zapier's 47.[9][11]
Each task seeds a fresh simulated company. Zapier's leaderboard page says the environment deliberately includes "traps: stale rows, near-duplicate names, policies buried in an inbox," and that the agent works alone after "one trigger message," with "no clarifying questions, no human in the loop."[3] Behind each app, Pydantic models are the source of truth for state. They mimic real API schemas, including pagination, required fields and common 4xx error responses, but everything runs locally so that runs are reproducible.[1][3]
Tool interface
In the default "API" mode used for the leaderboard, the agent has two tools. Search runs a BM25 keyword search over all the available API schemas and returns the top five matches. Execute mimics a curl or fetch request with a method, URL and body. No authentication is simulated.[1][3] Finding the right endpoints is part of the test: the paper notes that "the correct app or action name might not be known."[1] Agents get at most 50 steps per task. The paper says the limit is rarely reached and that several tools can be called in parallel within one step.[1] The repository also has a "Zapier" toolset built from Zapier's internal action schemas and a "Limited Zapier" toolset that exposes only the tools a task needs. These are for experiments and do not count toward official scores.[1][3][4]
Task construction and the Zapier-data claim
Zapier bases the benchmark's realism on the size of its own platform. The launch post says the domains were "selected based on the most common use-case patterns across the 3.7M companies and 2B monthly tasks Zapier sees," and elsewhere that "our platform processes over 2 billion AI tasks per month across 3.7 million total companies."[2] The leaderboard page says the benchmark is "built on real patterns from 2B+ monthly tasks across 3.7M companies."[3] These are Zapier's own figures and have not been independently verified.
The paper gives a narrower account. The tasks "were synthetically generated based on use cases from real customers," and "no PII or confidential customer data was used." What the authors did use was "the shape of workflows sent along with negative feedback on Zapier's Agents service." The authors say they targeted workflows that are hard to set up, clustered them by domain, and fed the patterns to models (Opus 4.6, GPT 5.3 Codex and Gemini 3) to generate tasks, then made "many passes on realism, difficulty, and variety of apps."[1] The disclaimer adds that "all data within tasks is entirely fictional" and that app behavior "does not represent exact production behavior."[1]
To make tasks harder, the authors added irrelevant data, hid key information behind tool-call responses, left it ambiguous where information could be found, gave incorrect entries similar names, and wrote strict business rules with overriding priorities. Tasks that were too easy were hardened further. The private set uses "additional hardening techniques to prevent overfitting on a small set of challenges."[1] Because some tasks tell the agent to follow policy documents that override external requests, which can look like prompt injection, task prompts explicitly point to those documents so that policy-following is not mixed up with obeying an injection.[1]
Every task must produce at least one concrete state change, be decidable by invariant-based checks, avoid free-form success criteria, record any judgment call as a structured state change, and run from a single instruction without further interaction.[1] To check that tasks are solvable, the authors wrote a hint sheet naming the endpoints to call (but not the parameter values). With the hints, even small models such as Haiku scored about 80 to 100%.[1]
Scoring
Only the final state of each connected app is graded. The paper says "how the Agent got there is not a concern," so an agent that updates a record five times instead of once scores the same, only at higher cost.[1] Tasks carry both positive assertions and negative ones. The negative assertions exist to stop "shotgun" reward hacking, such as emailing everyone in the company instead of the named recipient.[1][3] All criteria are "exact string matches and structural checks over the world state." The authors say this makes scores "largely reproducible across runs" and avoids one model grading another, while conceding that synthetically generated data "still carries some of this risk."[1]
The official metric, task_completed_correctly, is strict pass/fail: a task counts only if every scored assertion passes. A second metric, partial_credit, is the fraction of assertions passed. Zapier treats it as a diagnostic and a dense reward signal for training, not part of the headline score.[3][4] Zapier's page puts it this way: "Like production, mostly-right is still wrong."[3] Scores come from a single run. The paper and the leaderboard both say run-to-run variance is "typically within 1%."[1][3] The leaderboard also reports average cost per task in US dollars. The paper calls cost the main efficiency metric because timing is hard to compare across rate-limited providers.[1]
Example task
Zapier's leaderboard page walks through a sales task, sales.multi_hop_lookup. The agent is told to mark the "Meridian Corp Platform Deal" as won and route the win notice according to company policy. The CRM holds three near-duplicate Meridian opportunities. The spreadsheets have stale exchange-rate and tier rows, so the newest row must be used (€120,000 × 1.30 = $156,000, tier Enterprise). The routing policy exists only in an email, and an open critical support case sits on the parent account one level up. Six assertions check the outcome: three positive (the right opportunity closed, emails to the executive team and to support escalation) and three negative (no notices to the wrong teams). Passing five of the six gives a partial_credit of 0.83 but still fails the task.[3] The paper's appendix includes two more fully worked tasks: a meeting-conflict task with a time-zone decoy and an external email demanding a priority override, and an Asana task that must ignore emails marked "DRAFT" or "SUPERSEDED."[1]
Versions
| Version | Date | Changes |
|---|---|---|
| 0.1.0 | March 5, 2026 | Initial commit[7] |
| 0.1.1 | April 2, 2026 | Benchmark renamed AutomationBench; finance and HR domains added April 16-18[7] |
| (launch) | April 20-21, 2026 | Public release, blog post, arXiv paper, MIT license added[1][2][7] |
| 1.0.1 | June 6, 2026 | Fixed "some false positives and diversions from real apis"; license headers added[7] |
| 1.0.5 | July 16, 2026 (changelog) | Public baseline scores added to the README; tasks expanded and refined for fairness; better API fidelity and runner reliability[5] |
| 1.0.6 | July 31, 2026 (changelog; pushed to GitHub August 4) | Better task discoverability and fairness in public and private sets, including a Google Drive route to spreadsheet IDs for some Sheets tasks; less strict formatting checks; private tasks "made a bit harder to compensate for bug fixes"; Verifiers framework 0.2.0; prompt caching, refusal tracking and refreshed official runs[5][7] |
The README says private tasks "are sometimes made even harder in version updates to keep the benchmark around the same top score when fixing bugs," and that Zapier tries to rerun all models when a change moves scores beyond run-to-run variance.[4] This means scores from different versions are not directly comparable. Claude Opus 4.7 at max effort scored 9.9% in the April paper but is listed at 13.39% on the release 1.0.6 leaderboard.[1][3]
Results
Launch results
The paper's leaderboard, taken from the April 2026 release, had no model above 10%:[1]
| Model | Score | Cost per task |
|---|---|---|
| Claude Opus 4.7 (max) | 9.9% | $1.80 |
| Gemini 3.1 Pro (high) | 9.6% | $0.54 |
| GPT-5.4 (high) | 7.6% | $1.93 |
| Claude Sonnet 4.6 (max) | 5.3% | $1.81 |
| Claude Haiku 4.5 | 1.5% | $0.18 |
| GPT-5.4 (no reasoning) | 1.2% | $0.19 |
The authors reported that the top two models solved quite different tasks. The passing sets of Gemini and Opus had a Jaccard similarity of 0.17, and 71% of the tasks Opus solved were not solved by Gemini. In the Zapier toolset and the Limited Zapier toolset, Gemini 3.1 Pro's pass rate rose from 9.6% to 12.8% and 14.3%. On the unscored simple domain, even Haiku scored 97%.[1]
Zapier leaderboard (release 1.0.6)
Zapier's leaderboard at zapier.com/benchmarks lists 118 model and effort configurations on release 1.0.6, scored on the private set. The top rows, as accessed on September 30, 2026:[3]
| Rank | Model (effort) | Score | Cost per task |
|---|---|---|---|
| 1 | Claude Sonnet 5.5 (default fallbacks, max) | 44.75% | $1.14 |
| 2 | Claude Opus 5.5 (default fallbacks, max) | 42.47% | $1.44 |
| 3 | GPT-6 Astra (max) | 41.4% | $1.73 |
| 4 | GPT-6 Astra (xhigh) | 38.96% | $1.50 |
| 5 | GPT-6 Astra (high) | 37.14% | $1.44 |
| 6 | Claude Sonnet 5.5 (default fallbacks, xhigh) | 36.83% | $0.48 |
| 7 | Claude Opus 5.5 (default fallbacks, xhigh) | 35.77% | $0.89 |
| 8 | GPT-6 Astra (medium) | 34.09% | $1.27 |
| 9 | GPT-6 Sol (xhigh) | 33.2% | $0.27 |
| 10 | Claude Opus 5.5 (default fallbacks, high) | 33.03% | $0.71 |
| 12 | Claude Fable 5.1 (with Opus 5 fallback) | 31.4% | $2.45 |
| 14 | Gemini 3.7 Flash (high) | 30.44% | $0.61 |
| 20 | GPT-5.6 Sol (max) | 28.77% | $0.67 |
| 21 | Claude Opus 5 (max) | 26.94% | $3.05 |
| 31 | Kimi K3 | 22.68% | $0.43 |
| 32 | Claude Fable 5.1 (max) | 22.37% | $2.45 |
| 35 | GPT-6 Luna (max) | 20.7% | $0.04 |
| 41 | DeepSeek V4 Flash (max) | 18.11% | $0.04 |
| 112 | Claude Haiku 4.5 | 0.46% | $0.07 |
The footnotes change how some rows should be read. "Default fallbacks" means that a step refused by Anthropic's safety classifiers is sent to the provider's default fallback model. Zapier says refused tasks "were rerun with Anthropic's default fallback routing; other task results are unchanged." In the Claude Fable 5.1 row with an Opus 5 fallback, Opus 5 handled about 40% of tasks (260 of 657), and those completions count toward the 31.4%. The listed cost covers Fable 5.1 only, "so the true combo cost is higher." Gemini costs use standard list prices rather than promotional ones.[3] Zapier also publishes a per-domain view. On September 30, 2026, Operations had the highest top score (65.0%, Claude Opus 5.5 at max) and HR the lowest (31.67%, Claude Sonnet 5.5 at max).[3] GPT-6.1 Sol, which OpenAI released on September 29 with AutomationBench figures, was not yet on Zapier's leaderboard when it was accessed on September 30.[3][13]
Public-set baselines
After a user asked for reference numbers in July 2026, Zapier added public-set pass rates to the README with version 1.0.5 on July 16; that first table was led by Claude Opus 4.8 at 30.33%, with Gemini 3.5 Flash at 14.83%.[5][7][19] Zapier replaced it with re-run figures in the 1.0.6 release on August 4 and added a Claude Opus 5 row the same day.[4][7] At each model's highest reasoning effort, the current table lists Claude Opus 5 at 50.3%, Kimi K3 at 46.67%, Claude Fable 5 at 46.17%, GPT-5.6 Sol at 45.83%, Gemini 3.6 Flash at 45.00% and GLM 5.2 at 26.17%. It notes that the Fable 5 score comes after Anthropic's "July 2026 stricter classifier."[4] Public-set scores run well above private-set scores (Claude Opus 5 is at 26.94% on the private leaderboard), and Zapier says the ordering can differ because the private set uses "out of distribution complexity."[3][19]
Use in model launches
Several labs have put AutomationBench in their launch tables, and the setup behind each figure varies. The table below lists the reports this article checked against the primary source.
| Source | Model: score | Setup stated by the source |
|---|---|---|
| Anthropic, Claude Opus 5.5 launch (September 22, 2026)[15] | Opus 5.5: 40.0%; GPT-6 Astra: 41.4%; Fable 5.1: 31.4%; GPT-5.6 Sol: 28.8%; Opus 5: 26.9% | Run and reported by Zapier. Opus 5.5 was evaluated during early access "without fallback models, so safeguard interventions were considered failures." Other models are taken from Zapier's public leaderboard |
| Anthropic, Claude Sonnet 5.5 system card (September 28, 2026)[16] | Sonnet 5.5: 44.7%; Opus 5.5: 42.5%; GPT-6 Sol: 32.0%; Sonnet 5: 10.7% | Release 1.0.6 private set, max effort, default fallbacks. Opus 5.5's refused tasks rerun with fallbacks (40.0% without the rerun) |
| OpenAI, GPT-6 Sol and Luna launch (September 2026)[14] | GPT-6 Sol: 33.2% (xhigh, $0.27), 32.0% (max); GPT-6 Luna: 20.7% (max); Claude Opus 5: 26.9% (max, $3.05); Fable 5.1 w/ Opus 5 fallback: 31.4% | AutomationBench 1.0.6. OpenAI says GPT-6 Sol at xhigh "outperforms Claude Opus 5 at max effort at just 9% of Opus 5's cost per task" |
| OpenAI, GPT-6.1 Sol launch (September 29, 2026)[13] | GPT-6.1 Sol: 36.1% (max, $0.30), 35.5% (xhigh), 31.7% (medium, $0.19); Opus 5.5 w/ fallbacks: 42.5% (max, $1.44); GPT-6 Astra: 41.4% (max, $1.73); Fable 5.1 w/ Opus 5 fallback: 31.4% | Figures read from the embedded chart data, AutomationBench 1.0.6. The Opus 5.5 and Fable 5.1 points match Zapier's leaderboard |
| Google, Gemini 3.7 Flash launch (August 13, 2026)[17] | Gemini 3.7 Flash: 30.4%; Gemini 3.6 Flash: 17.0% | Not stated in the blog post. The figures match Zapier's leaderboard rows for the high setting |
| DeepSeek, DeepSeek-V4.1-Flash technical report (September 10, 2026)[18] | V4.1-Flash: 54.8 | "The public evaluation set of AutomationBench v1.0.6" with the benchmark's official scaffold. The comparison figures (Claude Opus 5 50.3, GPT-5.6 Sol 45.8, Kimi K3 46.7) are the same as Zapier's README public-set numbers |
Fallback accounting has become a recurring footnote. OpenAI's GPT-6.1 Sol post says: "The datapoint for Claude Fable 5.1 understates its actual cost, as it omits the cost of fallbacks, which occurred on ~40% of tasks." The GPT-6 Sol and Luna post carries almost the same note, naming "the Opus 5 fallbacks."[13][14] In the GPT-6.1 Sol post, OpenAI stresses that its new model scores 2.2 points above Opus 5.5 at medium effort "at roughly a third of the cost." The chart data also shows that Anthropic's Opus 5.5 configuration at max effort (42.5%) scored above both GPT-6.1 Sol (36.1%) and GPT-6 Astra (41.4%).[13] Anthropic, for its part, says the zero-fallback Opus 5.5 run that Zapier performed during early access "resulted in a lower score than Claude Opus 5.5 would achieve in practice."[15] DeepSeek's public-set figure of 54.8 is not comparable with private-set leaderboard scores. Neither are Artificial Analysis's AutomationBench-AA percentages, described below.[4][11][18] Reviewing the GPT-6 Astra launch, Vellum called AutomationBench "the biggest professional-work gap in the whole announcement" (41.4% for Astra against 31.4% for Fable 5.1).[20]
AutomationBench-AA
AutomationBench-AA is Artificial Analysis's independent run of the benchmark. Its methodology page describes it as "Artificial Analysis' run of Zapier's AutomationBench."[11] Artificial Analysis says it partnered with Zapier to run on "their private benchmark subset."[9] The main difference is the scoring. Artificial Analysis splits every AutomationBench assertion into an objective (something the agent must make true) or a guardrail (a check that already passes at the start and must not be broken). A task scores zero if any guardrail is violated or the task errors out. Otherwise it scores the share of objectives completed.[11] When it announced index version 4.3, the firm explained the name: "We call our implementation AutomationBench-AA because we award partial credit for completed objectives, with any guardrail violation reducing the task's score to zero." It also reports a separate "Tasks Completed" figure, the share of workflows where every objective is met with no violation, which is closer to Zapier's strict metric.[10]
| Zapier leaderboard | AutomationBench-AA | |
|---|---|---|
| Operator | Zapier | Artificial Analysis, with Zapier's collaboration[9] |
| Task set | Private held-out set[3] | Private 657-task held-out split, dataset version 1.0.6[11] |
| Headline metric | Share of tasks with every assertion passing[3] | Mean per-task share of objectives completed, zero on any guardrail violation[11] |
| Runs | Single run[1] | One run per task, 50-turn cap, API toolset[11] |
| Grader | Deterministic assertions[1] | Programmatic checks, "does not use a separate LLM judge"[11] |
The two scales are far apart. On September 7, 2026, GPT-6 Astra (max) scored 68.5% on AutomationBench-AA but completed every objective without a violation on only 41.6% of workflows. Artificial Analysis gave 32.1% on that measure for Claude Fable 5.1 (max with fallback) and 28.3% for Claude Opus 5 (max).[10] When AutomationBench-AA launched in July 2026, Claude Fable 5 (48.6%) and Claude Opus 4.8 (48.5%) led, and Artificial Analysis reported that Fable 5 "fell back to Opus on ~18% of tasks" under Anthropic's new classifier. The firm also found that "every model breaks business rules," with guardrail violations ranging from 0.46 per task (Gemini 3.5 Flash) to 1.26 (Qwen3.7 Plus). It found Finance the hardest domain, with agents completing around a third of Finance objectives against about 60% for Support and Operations.[9] Its launch article says models are scored against "nearly 12,000 assertions Zapier built."[9]
As of September 30, 2026, the AutomationBench-AA leaderboard was led by Claude Sonnet 5.5 (max effort, default fallback) at 71.3%, followed by Claude Opus 5.5 (max effort, default fallback) at 69.5%, DeepSeek V4.1 Flash (max) at 68.9% and GPT-6 Astra (max) at 68.5%. GPT-6.1 Sol scored 66.6% at xhigh and 64.9% at max.[12] AutomationBench-AA carries 5% of the Intelligence Index, alongside AA-Briefcase (15%) and GDPval-AA (10%) in the 30% Agents category. Because its test set is held out, Artificial Analysis said the share of the index made up of evaluations with private tasks or answers rose from 40% to 45% in version 4.3.[10][11]
Observed failure modes
The paper's main finding is that models often think they have succeeded when they have not. "More often than not, models declared success while actually failing": 72% of Opus's failures, 91% of Gemini's and 84% of GPT-5.4's involved this false confidence.[1] Other common failures were giving up when generic searches found nothing, assuming data lives in the CRM when it is actually in a spreadsheet, handling only part of a list (for example, some of twelve emails) before reporting the job done, and paraphrasing values that the instructions required verbatim.[1][3] The paper also found that Claude Opus 4.7 used an average of 12.6 steps and 29.8 tool calls per task, against 21.8 steps and 35.4 calls for Gemini 3.1 Pro.[1] Artificial Analysis saw similar differences in working style. GPT-5.5 (xhigh) averaged 49 tool calls over 25 turns, Claude Opus 4.8 (max) averaged 35 calls in 14 turns, and Grok 4.3 (high) took the fewest turns but scored lower, "consistent with declaring tasks complete prematurely."[9]
Limitations and criticism
The authors acknowledge the main weaknesses. Synthetic data risks "lack of realism and impossibility," the task set is too large to review fully by hand, and although they audited regularly for bugs and unrealistic complexity, "there is still room for improvements."[1] The README says that "there are still many we have not caught" and invites bug reports.[4] Because the private set is not released, outsiders cannot check the tasks behind leaderboard scores directly. Zapier says public and private results should show "directional agreement" but may not match "1:1."[4]
Public GitHub issues show the kinds of problems users have found:
- Grading and task defects. On September 29, 2026, an outside reviewer reported that a source-level audit at a pinned commit had flagged 108 of the 600 public tasks (18.0%) for corrections: 43 with missing or conflicting task information and 65 with grading defects. Examples included a task whose initial state lacked configuration IDs required by its assertions, and a finance task whose grader followed a stale instruction instead of a later correction. The issue was open and unanswered when accessed.[21] An earlier issue listed individual defects, such as a finance task that depended on a current date never given to the model, and a negative check that deleting a row could get around.[22]
- Toolset gaps. A September 23 issue reported that in the Limited Zapier toolset, all 100 HR tasks hide spreadsheet IDs but do not grant the Google Drive search tool needed to find them, which affects the 39 tasks that must change spreadsheet state. The same gap was reported in six Marketing tasks.[23] Version 1.0.6 had already added Google Drive access for some Sheets tasks, according to its changelog.[5]
- Prompt conflicts. A July issue pointed out that the global system prompt told agents to handle exclusions silently, while some tasks graded agents on explaining rejections. The maintainer replied that a fix was made in the 1.0.6 release.[24]
- Reproducibility. In July 2026 an independent team reran the public set and matched six of seven README scores to within 0.04 to 6.8 percentage points. Gemini 3.5 Flash was the exception: 29.9% in their run against 14.83% in the 1.0.5 README (the 1.0.6 README, published August 4, lists 38.33%). They suggested the gap might come from how the runner passes Gemini's thinking settings through an OpenAI-compatible endpoint, or from Gemini's thought signatures not being carried across turns. The maintainer confirmed that "we do use a proxy internally" and said they would check the scores; the issue remained open.[19] Another user asked in August why several open-weight models' scores rose sharply after an update, and that question had not been answered.[25]
The benchmark's claim to reflect real business work rests on Zapier's own description of its platform data. The paper says only the "shape" of workflows, drawn from negative feedback on Zapier Agents, informed synthetic tasks written with the help of frontier models from three of the labs whose models are now ranked.[1][2] Differences in configuration also make cross-lab comparisons fragile: whether Anthropic's safety-classifier fallbacks are on, whether a score comes from the public or private set, and whether it uses Zapier's strict metric or Artificial Analysis's partial-credit metric can each move a model's number by many points.[3][11][13][15][16]
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33 ^34 ^35 ^36 ^37 ^38 ^39 ^40 ^41 ^42Shepard, Daniel; Salimans, Robin. "AutomationBench." arXiv:2604.18934, submitted April 21, 2026. arxiv.org/...2604.18934
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Clifton, Anna Marie. "Introducing AutomationBench." Zapier Blog, April 20, 2026. zapier.com/...introducing-automationbench
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28Zapier. "AutomationBench: AI Agent Benchmarks" (leaderboard, release 1.0.6). Accessed September 30, 2026. zapier.com/benchmarks
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12Zapier. "AutomationBench" GitHub repository, README. Accessed September 30, 2026. github.com/...AutomationBench
- ^1 ^2 ^3 ^4 ^5Zapier. "Changelog" (CHANGELOG.md), AutomationBench repository. Accessed September 30, 2026. github.com/...CHANGELOG.md
- ^Zapier. "LICENSE," AutomationBench repository. github.com/...LICENSE
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Zapier. Commit history and pyproject.toml versions, AutomationBench repository. Accessed September 30, 2026. github.com/...main
- ^Prime Intellect (@PrimeIntellect). Post on X announcing AutomationBench on the Environments Hub, April 21, 2026. x.com/...2046393268842451187
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Artificial Analysis. "Announcing AutomationBench-AA." July 6, 2026. artificialanalysis.ai/...zapier-automationbench-aa
- ^1 ^2 ^3 ^4 ^5Artificial Analysis. "Announcing the Artificial Analysis Intelligence Index v4.3." September 7, 2026. artificialanalysis.ai/...s-intelligence-index-v4-3
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13Artificial Analysis. "Artificial Analysis Intelligence Benchmarking Methodology" (AutomationBench-AA section and version history). Accessed September 30, 2026. artificialanalysis.ai/...intelligence-benchmarking
- ^Artificial Analysis. "AutomationBench-AA: Agentic SaaS Workflow Benchmark." Accessed September 30, 2026. artificialanalysis.ai/...automationbench-aa
- ^1 ^2 ^3 ^4 ^5 ^6OpenAI. "Introducing GPT-6.1 Sol." September 29, 2026. openai.com/...introducing-gpt-6-1-sol
- ^1 ^2OpenAI. "Introducing GPT-6 Sol and Luna." September 2026. openai.com/...introducing-gpt-6-sol-and-luna
- ^1 ^2 ^3 ^4Anthropic. "Introducing Claude Opus 5.5." September 22, 2026. anthropic.com/claude-opus-5-5
- ^1 ^2 ^3Anthropic. "System Card: Claude Sonnet 5.5." September 28, 2026. anthropic.com/claude-sonnet-5-5-system-card
- ^1 ^2Google. "Gemini 3.7 Flash: our most intelligent workhorse model." August 13, 2026. blog.google/...introducing-gemini-3-7-flash
- ^1 ^2 ^3DeepSeek-AI. "DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression" (technical report). September 10, 2026. huggingface.co/...DeepSeek_V41_Tech_Report.pdf
- ^1 ^2 ^3zapier/AutomationBench issues #5 ("Request: Release public-set baseline results") and #7 ("Discrepancy of gemini-3.5-flash public-set pass rate: independent rerun measures 29.9% vs README's 14.83%"). July 2026. github.com/...issues
- ^Zeeb, Nicolas. "GPT-6 Astra Benchmarks Explained." Vellum, September 3, 2026. vellum.ai/...gpt-6-astra-benchmarks-explained
- ^zapier/AutomationBench issue #28, "Data quality issues affecting 108 cases." September 29, 2026. github.com/...28
- ^zapier/AutomationBench issue #24, "Bugs with benchmark tasks." September 8, 2026. github.com/...24
- ^zapier/AutomationBench issue #27, "[limited_zapier] HR tasks hide Sheets' spreadsheet IDs and names but omit the Drive tool needed to find them (affecting 39/100 tasks)." September 23, 2026. github.com/...27
- ^zapier/AutomationBench issue #11, "System Prompt vs User Prompt conflict: silent exclusion rule contradicts tasks that require explaining rejections." July 29, 2026. github.com/...11
- ^zapier/AutomationBench issue #17, "Question about the significant performance boost in the latest update." August 20, 2026. github.com/...17
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,843 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent verification V5 (xg14, 30 Sep 2026): ~115 claims vs arXiv paper, Zapier blog/board (118 rows), GitHub repo/changelog/issues, AA, OpenAI charts; 1 material (README version) + 5 minor fixed
Cite this page: AI Wiki. "AutomationBench." aiwiki.ai, updated 30 Sept 2026, fact-checked 30 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/automationbench