# FrontierCode

> Source: https://aiwiki.ai/wiki/frontiercode
> Updated: 2026-09-29
> Fact-checked: 2026-09-29
> Categories: AI Agents, AI Benchmarks, AI Code Generation, Model Evaluation
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "FrontierCode." aiwiki.ai, 29 Sept 2026. https://aiwiki.ai/wiki/frontiercode
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

**FrontierCode** is an agentic coding benchmark created by [Cognition](https://aiwiki.ai/wiki/cognition_ai), the company behind the [Devin](https://aiwiki.ai/wiki/devin) coding agent. Instead of asking only whether an AI agent's patch passes unit tests, it asks whether a repository's maintainer would actually merge the change, grading correctness together with test quality, scope discipline, style and adherence to codebase conventions.[1] Cognition introduced the benchmark on 8 June 2026 with 150 tasks written by open-source maintainers from 36 repositories, and released a revised methodology, FrontierCode 1.1, on 7 July 2026.[1][2] Cognition runs the evaluations, keeps the tasks private and publishes results on a public leaderboard.[1][3] By September 2026 FrontierCode results appeared in model launch materials from [Anthropic](https://aiwiki.ai/wiki/anthropic), [OpenAI](https://aiwiki.ai/wiki/openai), Google DeepMind and SpaceXAI, and Anthropic's launch post for Claude Sonnet 5.5 (28 September 2026) described it as a test of "whether a code change could be merged without human edits."[5][15][18][19]

## Overview

| Attribute | Detail |
| --- | --- |
| Developer | Cognition[1] |
| Introduced | 8 June 2026 (FrontierCode 1.0)[1] |
| Current version | FrontierCode 1.1, released 7 July 2026[2][3] |
| Tasks | 150, from 36 open-source repositories[1] |
| Subsets | Extended (all 150 tasks), Main (100 hardest), Diamond (50 hardest; deprecated in 1.1)[1][2] |
| Task authors | More than 20 open-source maintainers, each spending more than 40 hours per task[1] |
| Metrics | Score (weighted rubric aggregate, zero if any blocker fails) and pass rate (clears all blockers)[1] |
| Trials | 5 runs per model at every available reasoning effort; best effort reported[1] |
| Task availability | Not released publicly, to avoid contamination[1] |
| Leaderboard | cognition.com/frontiercode, launched 17 July 2026[3] |

## History

Cognition published "Introducing FrontierCode" on 8 June 2026, with a byline crediting Eric Lu, Ben Pan, Deniz Birlikci, Sam Lee, Ray Wang, Rohan Choudhury, Fermi Ma, TC Qin, Carlo Baronio, Silas Alberti and others.[1] The company's announcement post on X that day said "Models write sloppy code that works but isn't maintainable. Our eval is first to measure: would you actually merge this code?"[22] Cognition's chief executive [Scott Wu](https://aiwiki.ai/wiki/scott_wu) framed it as a response to SWE-bench-style grading, writing that "passing a unit test is only one part of writing production-ready code."[23]

The blog post cites a 10 March 2026 research note from [METR](https://aiwiki.ai/wiki/metr), "Many SWE-bench-Passing PRs Would Not Be Merged into Main", as motivation.[1] METR's note found that "roughly half of test-passing SWE-bench Verified PRs written by mid-2024 to mid/late-2025 agents would not be merged into main by repo maintainers."[28] The [Latent Space](https://aiwiki.ai/wiki/latent_space) newsletter wrote that FrontierCode "was explicitly inspired and named for FrontierMath", referring to [Epoch AI](https://aiwiki.ai/wiki/epoch_ai)'s mathematics benchmark [FrontierMath](https://aiwiki.ai/wiki/frontiermath).[27]

| Date | Event |
| --- | --- |
| 8 June 2026 | FrontierCode 1.0 introduced with Extended, Main and Diamond subsets[1] |
| 9 June 2026 | Anthropic's Claude Fable 5 launch reports FrontierCode 1.0 results[9][10] |
| 7 July 2026 | FrontierCode 1.1: fair-internet-use rules, 75 blockers relaxed, Diamond deprecated[2] |
| 17 July 2026 | Public leaderboard launched with 1.0 and 1.1 results[3] |
| 6 August 2026 | Leaderboard adds a flag-rate column for unfair internet use[3] |
| 10 September 2026 | Cognition's own SWE-2 model added[3][17] |
| 22 September 2026 | Claude Opus 5.5, GPT-6 Sol and GPT-6 Luna added[3] |
| 28 September 2026 | Claude Sonnet 5.5 added[3] |

## Task design

Cognition says it worked directly with the maintainers of 36 "flagship" open-source repositories, and that more than 20 of them built the tasks from the repositories they maintain, each spending more than 40 hours per task through several rounds of iteration with other evaluation engineers and Cognition researchers.[1] Maintainers quoted in the launch post included Tomer Nosrati of Celery, Martin McKeaveney of Budibase, Merlijn Vos of Uppy and Claudio Costa of Mattermost; Nosrati's summary was "Where others grade like a CI, FrontierCode grades like a tech lead."[1]

According to Cognition, other benchmarks built their issues from single pull requests by programmatic scraping, while FrontierCode tasks are hand-selected by maintainers from multi-PR chains and freeform requests, and cover three times as many programming languages as [SWE-Bench Pro](https://aiwiki.ai/wiki/swe_bench_pro).[1] The FrontierCode 1.1 post adds that tasks are "sourced from real PRs in open-source codebases."[2] Anthropic's system cards give examples: fixing websocket bugs in aiohttp, hardening Prisma's browser bundle, and extending JSON schema linting rules.[6][10]

Each prompt has two parts: a task description, and codebase guidelines for testing, linting and style "just like those found in AGENTS.md." Cognition describes the task descriptions as "humanlike and deliberately concise", a third the length of SWE-Bench Pro's, and says the agent is expected to infer the maintainer's intent from the same context a human contributor would have.[1] Difficulty is scaled through quality rubrics rather than larger patches; Cognition says FrontierCode is harder for agents than DeepSWE despite having smaller patches.[1] Anthropic describes the agent as working autonomously in a container "with internet access" and "no timeout information."[6][8]

The published example task comes from the C++ jsonschema repository and asks the agent to route every warning message through a new `LOG_WARNING()` helper. Cognition reports that Claude Opus 4.8 repeatedly called the helper for the first line of a multi-line warning and wrote the remaining lines straight to `std::cerr`. The output was the same, but the code assumed the helper and `std::cerr` would always be the same stream, and a blocking criterion failed it.[1]

Cognition does not plan to release the tasks publicly "to avoid contamination", but says it is opening the evaluation to all model creators.[1]

## Subsets

FrontierCode 1.0 defined three nested subsets of increasing difficulty: Extended (all 150 tasks), Main (the 100 hardest, including Diamond) and Diamond (the 50 hardest).[1] With FrontierCode 1.1, Cognition stopped reporting Diamond, saying the revised methodology meant Diamond "no longer reflects the 50 hardest tasks" and that, because solve rates on the hardest tasks are so low, "Diamond performance is inherently noisy."[2] Since July 2026 results are reported on Main and Extended only.[2][3]

## Grading

Cognition grades each patch along six axes: behavioral correctness, regression safety, mechanical cleanliness (build, lint and style checks), test correctness, scope, and code quality.[1] It combines classical and new methods:

| Criterion type | Method | Passes when |
| --- | --- | --- |
| Behavioral correctness | Classical: injects test files, runs them, cleans up | All injected tests pass |
| Mechanical cleanliness, regression safety | Command: runs a shell command | Exit code 0 |
| Test correctness | Reverse-classical: runs the agent's own tests against the base commit | The tests fail |
| Behavioral correctness for complex tasks | Adaptive classical grading: an LLM adapts reference tests or application code to the agent's implementation | Adapted tests pass |
| Scope | Checks allowed and denied files, diff-size limits and, optionally, semantic locality | Diff within constraints |
| Code quality | Prompt: an LLM reviews the diff against a natural-language prompt | LLM score meets threshold |

Source: Cognition, "Introducing FrontierCode".[1]

The reverse-classical check requires that tests the agent writes fail on the original, broken code, which Cognition presents as an automatic way to confirm the agent understood the problem. The scope criterion enforces restraint through file rules, size limits (changed lines, net growth, files touched) and LLM-based checks of where inside a file a change lands. Adaptive classical grading uses an internal tool called mutagent, which uses an LLM to patch the test environment or application code so that valid solutions with different function names or error wording are not failed on superficial differences.[1] The code-quality and adaptive methods make FrontierCode partly an [LLM-as-a-judge](https://aiwiki.ai/wiki/llm_as_a_judge) benchmark.

Each criterion is either a **blocker**, a requirement a maintainer would treat as a hard stop in code review (correctness, and also concerns such as performance or scope), or a **non-blocker** quality signal such as style, type safety or readability.[1] A solution that clears every blocker passes; its score is the weighted aggregate of the rubric items it satisfies. A solution that fails any blocker scores zero.[1] Cognition therefore reports two numbers: pass rate and score. Every model is run five times at each available reasoning effort, the metric is averaged across the five trials, and each model is reported at its best-performing effort level.[1]

### Quality control

Because rubric grading is subjective, Cognition built a review pipeline in which each task author first documents the rationale for every rubric item, then writes a "hack report": a deliberately lazy or adversarial solution to expose false positives, and a valid alternative solution to expose false negatives. Devin is also asked to find new ways to game the rubric. Authors must write four solutions targeting scores from 0 to 100% to calibrate the rubric. Tasks then pass through a pod lead and a final review by a Cognition researcher, who solves a random subset of tasks personally.[1]

Cognition claims an "81% lower false positive rate compared to SWE-Bench Pro", and elsewhere in the same post says FrontierCode "produces 81% less misclassification errors than other leading benchmarks", based on its analysis of agent trajectories.[1] These are Cognition's own measurements.

## FrontierCode 1.1

FrontierCode 1.1, published on 7 July 2026, made three changes.[2]

**Internet use.** Because tasks come from real pull requests, their solutions may exist in a later upstream version, on mirrors, or in package registries (an agent can sometimes get the fix by installing the latest release). Cognition says it saw a few such cases when building 1.0, but that newer models such as Claude Fable 5 were getting better at finding them.[2] It chose not to disable internet access, because some tasks require it (for example, looking up API contracts) and because models increasingly rely on search. It also rejected domain blocklists (its list grew to about 1,200 domains while agents kept finding workarounds) and allowlists.[2] Instead, 1.1 adds a prompt that defines fair internet use, such as reading documentation, and a classical verifier that flags references to source pull requests, upstream patches or solution-bearing mirrors; flagged runs score zero.[2] Cognition reports that with the prompt in place, unfair internet use fell below 1% for every model it evaluated.[2]

**Relaxed blockers.** Cognition audited more than 1,000 criteria and demoted 75 overly strict blockers to non-blockers, which it expects to reduce false negatives.[2]

**Diamond deprecated.** Main and Extended became the only reported subsets, as described under Subsets above.[2]

Cognition said that 1.1 changed absolute scores but that "the relative performances of the models we evaluated did not substantially change compared to 1.0."[2] Scores did rise substantially: Claude Opus 4.8's best Main score went from 34.3% under 1.0 to 46.5% under 1.1, and GPT-5.5's from 25.5% to 43.0%.[4]

## Leaderboard

Cognition's leaderboard, launched on 17 July 2026, shows FrontierCode 1.1 (the current revision) and 1.0 results on the Main and Extended subsets, with columns for score, pass rate, flag rate (the share of runs zeroed for unfair internet use), mean cost per rollout in US dollars and output tokens.[3] Each row shows a model's best-scoring reasoning effort.[3] The underlying data file labels the harness used for each model: Claude models run in claude-code, OpenAI models in codex, SpaceXAI's Grok models in grok-build, Cognition's SWE-2 in devin, Cursor's Composer 2.5 in cursor-cli, several open-weight models in mini-swe-agent, and SWE-1.7, SWE-1.6 and a number of other models in a harness labelled chisel.[4]

### FrontierCode 1.1 results (as of 29 September 2026)

The top of the FrontierCode 1.1 leaderboard, ordered by Main score. Each figure is the model's best reasoning effort on that subset; cost is mean US dollars per rollout on Main.

| Rank (Main) | Model | Main score (effort) | Main pass rate | Main cost | Extended score (effort) |
| --- | --- | --- | --- | --- | --- |
| 1 | [Claude Opus 5.5](https://aiwiki.ai/wiki/claude_opus_5_5) | 54.6% (medium) | 59.6% | $0.80 | 65.3% (medium) |
| 2 | [Claude Fable 5](https://aiwiki.ai/wiki/claude_fable_5) | 53.5% (xhigh) | 58.9% | $13.09 | 64.9% (xhigh) |
| 3 | [Claude Opus 5](https://aiwiki.ai/wiki/claude_opus_5) | 53.4% (medium) | 58.9% | $4.31 | 63.6% (medium) |
| 4 | [GPT-6 Astra](https://aiwiki.ai/wiki/gpt_6_astra) | 53.3% (max) | 58.8% | $4.59 | 64.5% (max) |
| 5 | [Claude Sonnet 5.5](https://aiwiki.ai/wiki/claude_sonnet_5_5) | 52.1% (xhigh) | 57.2% | $1.59 | 64.4% (xhigh) |
| 6 | [Claude Fable 5.1](https://aiwiki.ai/wiki/claude_fable_5_1) | 50.9% (medium) | 55.5% | $3.28 | 63.6% (medium) |
| 7 | SWE-2 (Cognition) | 50.0% (max) | 55.5% | $1.18 | 62.5% (max) |
| 8 | [GPT-6 Sol](https://aiwiki.ai/wiki/gpt_6_sol) | 49.3% (max) | 54.3% | $2.07 | 60.7% (max) |
| 9 | [Grok 4.6](https://aiwiki.ai/wiki/grok_4_6) | 48.0% (high) | 53.1% | $2.88 | 61.3% (high) |
| 10 | [Grok 4.7](https://aiwiki.ai/wiki/grok_4_7) | 47.6% (high) | 53.1% | $6.65 | 59.4% (high) |
| 11 | GPT-5.6 Sol | 47.5% (max) | 52.9% | $5.19 | 60.6% (max) |
| 12 | [Claude Opus 4.8](https://aiwiki.ai/wiki/claude_opus_4_8) | 46.5% (max) | 51.6% | $9.62 | 59.6% (max) |
| 13 | [Kimi K3](https://aiwiki.ai/wiki/kimi_k3) | 44.2% | 48.9% | $3.82 | 58.2% |
| 14 | [Gemini 3.7 Flash](https://aiwiki.ai/wiki/gemini_3_7_flash) | 43.6% (medium) | 48.9% | $1.82 | 56.3% (medium) |
| 15 | GPT-5.5 | 43.0% (xhigh) | 48.2% | $4.03 | 56.7% (xhigh) |
| 16 | [Claude Sonnet 5](https://aiwiki.ai/wiki/claude_sonnet_5) | 42.7% (xhigh) | 47.6% | $10.07 | 56.2% (xhigh) |
| 17 | Grok 4.5 | 42.4% (high) | 47.2% | $1.30 | 56.5% (high) |
| 18 | GPT-6 Luna | 42.4% (max) | 47.8% | $0.10 | 56.1% (max) |

Source: Cognition leaderboard and its data file, retrieved 29 September 2026.[3][4] The leaderboard listed 41 models on that date; the lowest Main scores were SWE-1.6 (9.4%) and Mistral 3.5 Medium (8.0%).[3] Flag rates were near zero for most models, with the highest on Main at DeepSeek V4 Flash 0731 (25.5% of runs) and DeepSeek V4 Pro 0813 (10.6%).[3]

### FrontierCode 1.0 results

At launch, Cognition reported that on Diamond "the best performing model, Claude Opus 4.8, achieves a score of only 13.4%", with GPT-5.5 at 6.3%, Gemini 3.1 Pro at 4.7% and Kimi K2.6, the best open-weight model, at 3.8%. It added that GPT-5.5 "consistently uses up to 4x fewer tokens than Opus 4.8."[1] A day later Anthropic reported Claude Fable 5 at 29.3% on Diamond.[10]

| Model | Diamond score | Main score | Extended score | Source |
| --- | --- | --- | --- | --- |
| Claude Fable 5 (xhigh) | 29.3% (pass rate 30.2%) | 46.3% (pass rate 48.8%) | not given | Anthropic Fable 5 system card[10] |
| Claude Fable 5 (xhigh) | not in data file | 46.8% | 61.6% | Cognition 1.0 leaderboard data[4] |
| Claude Opus 4.8 | 13.4% | 34.3% | 51.8% | Cognition launch post[1] |
| GPT-5.5 | 6.3% | 25.5% (xhigh) | 44.8% (high) | Launch post (Diamond); 1.0 data (Main, Extended)[1][4] |
| Gemini 3.1 Pro | 4.7% | 16.7% (high) | 34.2% (low) | Launch post (Diamond); 1.0 data[1][4] |
| Kimi K2.6 | 3.8% | 16% | 37% | Cognition launch post[1] |
| Claude Sonnet 4.6 | not given | 15.1% (high) | 33.6% (high) | Cognition 1.0 leaderboard data[4] |

The two Fable 5 Main figures differ by half a point (46.3% in Anthropic's system card, 46.8% in Cognition's current 1.0 data). The leaderboard changelog notes that Cognition corrected Fable 5's costs on the 1.0 leaderboard on 17 July 2026 without mentioning a score change, although the FrontierCode 1.1 post of 7 July said Cognition was releasing "updated scores for Fable 5."[3][2][10]

## Use in model launches

FrontierCode results are produced by Cognition, and model developers quote them in their launch materials. Anthropic's system cards say "Cognition ran the evaluation for every model shown and reported the results; Claude models were run in Claude Code and GPT models were run in Codex CLI."[6][8] Labs do not all report the same number. Anthropic's Opus 5.5 launch table uses max effort unless otherwise noted, and the Sonnet 5.5 table marks its effort levels, while Cognition's leaderboard and Anthropic's system-card prose use each model's best effort.[5][7][8]

| Launch (date) | Developer | Figure reported | Set and version | Source |
| --- | --- | --- | --- | --- |
| Claude Fable 5 (9 June 2026) | Anthropic | 29.3% Diamond, 46.3% Main (xhigh); launch post says Fable 5 "scores highest among frontier models, even at medium effort" | 1.0 | [9][10] |
| Claude Sonnet 5 (30 June 2026) | Anthropic | 38.8% (Sonnet 4.6 15.1%, GPT-5.5 25.5%) | "FrontierCode v1" | [11] |
| Claude Opus 5 (24 July 2026) | Anthropic | 53.4% Main, 63.6% Extended, both at medium (best effort); launch-post quote from Scott Wu: Opus 5 "approaches Fable-level performance at half the cost" | 1.1 | [12][13] |
| Grok 4.6 (12 August 2026) | SpaceXAI | 61.3% Extended (High), against Grok 4.5 High 56.6%, GPT-5.6 Sol Max 60.6%, Fable 5 Max 63.6% | 1.1 Extended | [19] |
| Gemini 3.7 Flash (August 2026) | Google DeepMind | 43.6% Main (Gemini 3.6 Flash 34.4%, Claude Sonnet 5 42.7%, GPT-5.6 Terra 41.3%) | 1.1 Main | [18] |
| Claude Fable 5.1 (1 September 2026) | Anthropic | 50.9% Main and 63.6% Extended, at medium | 1.1 | [14] |
| GPT-6 Astra (3 September 2026) | OpenAI | 53.3% Main, 64.5% Extended | 1.1 | [15] |
| SWE-2 (10 September 2026) | Cognition | 50.0% Main | 1.1 Main | [17] |
| Claude Opus 5.5 (22 September 2026) | Anthropic | 54.4% Main at max in the launch table; 54.6% Main and 65.3% Extended at medium (best) | 1.1 | [7][8] |
| GPT-6 Sol (22 September 2026) | OpenAI | No figure in the post's text; says Sol "is able to match Claude Fable 5.1 xhigh at much lower cost" | 1.1 | [16] |
| Claude Sonnet 5.5 (28 September 2026) | Anthropic | 46.2% Main at Max and 52.1% at Xhigh in the launch table; 64.4% Extended at xhigh, 59.1% at max | 1.1 Main and Extended | [5][6] |

Figures for the same model sometimes differ between documents. The Claude Opus 5.5 system card gives Claude Fable 5.1's best Main score as 52.8%, while the Fable 5.1 system card, OpenAI's GPT-6 Astra table and Cognition's leaderboard all give 50.9%.[8][14][15][3] The Claude Sonnet 5 system card's "FrontierCode v1" row does not name a subset. Its comparison figures for Sonnet 4.6 (15.1%) and GPT-5.5 (25.5%) match those models' best FrontierCode 1.0 Main scores in Cognition's data file.[4][11] The Opus 5 system card describes FrontierCode as "a coding evaluation run and scored by Cognition."[13] For its GPT-6 Astra runs, OpenAI disclosed that Astra "was run with a developer message similar to a section of its developer message in Codex", which asks the model to avoid excessive test files and "unrelated cleanup" and ends "The goal is clean, mergeable code." OpenAI added: "The prompt was not optimized for the eval."[15]

Cognition also published its own posts for Anthropic launches. For Opus 5.5 it called FrontierCode "our proprietary benchmark" and said Opus 5.5 "is the new #1", costing "less than a tenth of Fable 5 and about a sixth of Astra per task."[20] For Sonnet 5.5 it reported 64.4% on Extended, "surpassing even Fable 5.1 at extra high reasoning effort."[21]

### Claude Sonnet 5.5

Anthropic's Sonnet 5.5 launch table lists FrontierCode 1.1 (Main) at 46.2% at Max effort and 52.1% at Xhigh, against 42.4% for Claude Sonnet 5, 54.4% for Claude Opus 5.5 and 49.3% for GPT-6 Sol.[5] The footnote explains why Max is lower:

> Sonnet 5.5 scores lower at Max effort than at Xhigh. FrontierCode evaluates whether a code change could be merged without human edits. It penalizes out-of-scope changes, even if they are high-quality or helpful. At Max effort, Sonnet 5.5 more often ran Claude Code's code-review skill, which splits the review across many subagents, and in two cases Cognition examined, this led to a timeout or to extra edits beyond the task's scope, and therefore to a lower score.[5]

The post also says that at High effort, the Claude Platform default, Sonnet 5.5 "matches GPT-6 Sol's best score for about a fifth of the cost per task", and that at High it scores 10 points more than Sonnet 5 at the same setting "at about one fifteenth of the cost per task."[5] The Sonnet 5.5 system card says Anthropic reports the scores "without any changes to the official evaluation."[6] Cognition's data file gives the per-effort results:

| Effort | Main score | Main pass rate | Main cost per rollout | Mean output tokens (Main) | Extended score | Main flag rate |
| --- | --- | --- | --- | --- | --- | --- |
| Low | 29.3% | 32.8% | $0.19 | 6,353 | 43.4% | 0% |
| Medium | 36.5% | 40.6% | $0.24 | 7,597 | 50.5% | 0% |
| High | 49.4% | 54.4% | $0.42 | 13,003 | 61.5% | 0% |
| Xhigh | 52.1% | 57.2% | $1.59 | 47,385 | 64.4% | 0% |
| Max | 46.2% | 51.1% | $20.78 | 595,651 | 59.1% | 4.65% |

Source: Cognition leaderboard data file, retrieved 29 September 2026.[4] The same file gives Sonnet 5 at High 39.4% on Main at $6.10 per rollout, consistent with Anthropic's "10 points" and "one fifteenth" comparison, and GPT-6 Sol's best Main score as 49.3% at max for $2.07.[4]

## Effort levels and the scope penalty

Several developers have reported that FrontierCode scores do not rise steadily with reasoning effort, and attribute this to the scope and blocker rules:

- **Claude Fable 5.1.** Anthropic's system card says Fable 5.1's score peaks at medium effort because at higher efforts it "occasionally adds more small, unrequested changes in files outside the task, such as a documentation comment in an adjacent file, an edit to a docs page, or a new CI job where an existing one could have been reused." Adding a brevity instruction reduced such edits, but Anthropic reported scores without changing the official evaluation.[14]
- **Claude Opus 5.5.** The system card notes "a decline in model performance on FrontierCode (Main and Extended) above medium effort" and says scores "mostly recover at max."[8]
- **Claude Sonnet 5.5.** Lower at Max than at Xhigh, as described above.[5]

Cognition gave a general explanation in June 2026, replying to the YouTuber and developer Theo (@theo): because grading has "many blocking criteria beyond unit tests", small behavioral changes across reasoning efforts are amplified. In its example, a model at medium effort might add an unnecessary defensive patch that a lower-effort run never attempts and a higher-effort run correctly leaves out, so "small shifts in reasoning effort can move results in non-monotonic ways."[26]

## Reception and criticism

Latent Space's AINews issue of 9 June 2026 made FrontierCode its lead story and said the benchmark was "reshaping the discourse around what 'good coding performance' should mean."[27] The issue opened by noting that "it is rare that we are personally involved in the title story of the day."[27] Latent Space's author, swyx, lists affiliations with both Cognition and Latent Space on his X profile, and Cognition's acknowledgments credit him among FrontierCode's "Outstanding External Contributors."[1][24][27] On X, swyx described FrontierCode as the benchmark for a third era of AI coding, "Maintainable Code", after HumanEval-style autocomplete and SWE-bench-style test passing.[24]

The main published criticism concerns reproducibility. On launch day Theo wrote that the results chart "scares me a bit. My guess is that reproducibility is low", pointing out that in Cognition's chart Opus scored higher at low effort than at medium and higher at xhigh than at max.[25] He asked Cognition which repositories make up Diamond, how many runs it made per model, and how much scores vary on a rerun.[25] Cognition replied that it runs each model five times per effort, "Can't comment on the exact repo list", and would "aim to run N>>5 as a follow-up to quantify variance better."[26] One month later Cognition deprecated Diamond, in part because "Diamond performance is inherently noisy."[2]

Other structural limits follow from the design and are stated in Cognition's own materials:

- **Closed tasks.** The tasks are private and Cognition runs the evaluations, so outside parties cannot reproduce scores or check the rubrics directly.[1][3] Anthropic's system cards note that Cognition's result files "carry no uncertainty intervals", although Cognition's published figures show error bars.[6][8]
- **Owner also competes.** Cognition evaluates its own coding models (SWE-1.6, SWE-1.7 and SWE-2) on the benchmark, and its SWE-2 launch post leads with its FrontierCode score.[3][17]
- **Harness dependence.** Models are run in different agent harnesses (Claude Code, Codex, Grok Build, Devin and others), so a score reflects the model and its harness together.[4] The Sonnet 5.5 result at Max effort depended partly on a Claude Code skill that spawned subagents.[5]
- **Model-graded criteria.** Code-quality criteria and adaptive grading rely on LLM judgments, which Cognition controls through its rubric calibration process.[1]
- **Prompting.** OpenAI's disclosure that GPT-6 Astra ran with a developer message asking for "clean, mergeable code" shows that scores can depend on instructions beyond the task prompt.[15]

## Naming and related benchmarks

FrontierCode is distinct from several similarly named 2026 benchmarks. In the Claude Fable 5 launch post, a customer quote from Scott Wu said Fable 5 "is the highest-scoring model on FrontierBench, Cognition's frontier coding eval", while Anthropic's own text on the same page calls Cognition's evaluation FrontierCode.[9] FrontierBench v0.1 appears as a separate row from FrontierCode 1.1 in the Claude Opus 5 system card's evaluation table, which describes it as "a successor to Terminal-Bench 2.1 developed by the same team."[13] FrontierSWE is a separate benchmark of ultra-long-horizon engineering tasks from Proximal, reported in its own system-card section.[10][14] FrontierMath is Epoch AI's mathematics benchmark, which Latent Space describes as FrontierCode's naming inspiration.[27]

Among coding benchmarks, FrontierCode positions itself against [SWE-bench Verified](https://aiwiki.ai/wiki/swe-bench_verified) and SWE-Bench Pro, which Cognition says test functional correctness rather than quality and are prone to misclassification.[1] Other 2026 benchmarks reported next to it in launch tables include DeepSWE v1.1, [Terminal-Bench](https://aiwiki.ai/wiki/terminal_bench) and [CursorBench](https://aiwiki.ai/wiki/cursorbench).[5][15]

## References

1. Eric Lu, Ben Pan, Deniz Birlikci, Sam Lee, Ray Wang, Rohan Choudhury, Fermi Ma, TC Qin, Carlo Baronio, Silas Alberti et al. "Introducing FrontierCode." Cognition, June 8, 2026. https://cognition.com/blog/frontier-code
2. Eric Lu, Ben Pan, Fermi Ma, Alex Lombardi et al. "FrontierCode 1.1." Cognition, July 7, 2026. https://cognition.com/blog/frontier-code-1.1
3. Cognition. "FrontierCode Leaderboard" (including methodology revisions and changelog). Retrieved September 29, 2026. https://cognition.com/frontiercode
4. Cognition. FrontierCode leaderboard data file (FrontierCode 1.0 and 1.1 results by model and reasoning effort). Retrieved September 29, 2026. https://cognition.com/data/frontiercode-leaderboard/data.json
5. Anthropic. "Introducing Claude Sonnet 5.5." September 28, 2026. https://www.anthropic.com/claude-sonnet-5-5
6. Anthropic. "Claude Sonnet 5.5 System Card," section 8.4. September 28, 2026. https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c857e3250bec1/Claude%20Sonnet%205.5%20System%20Card.pdf
7. Anthropic. "Introducing Claude Opus 5.5." September 22, 2026. https://www.anthropic.com/claude-opus-5-5
8. Anthropic. "Claude Opus 5.5 System Card," section 8.4. September 22, 2026. https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf
9. Anthropic. "Claude Fable 5 and Claude Mythos 5." June 9, 2026. https://www.anthropic.com/news/claude-fable-5-mythos-5
10. Anthropic. "System Card: Claude Fable 5 & Claude Mythos 5," sections 8.1 and 8.4. June 9, 2026. https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf
11. Anthropic. "Claude Sonnet 5 System Card," sections 8.1 and 8.4. June 30, 2026. https://www-cdn.anthropic.com/283ef97c476cf442c91d9a37d5b214242a55bb92/Claude%20Sonnet%205%20System%20Card.pdf
12. Anthropic. "Introducing Claude Opus 5." July 24, 2026. https://www.anthropic.com/news/claude-opus-5
13. Anthropic. "System Card: Claude Opus 5," sections 8.1 and 8.4. July 24, 2026. https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf
14. Anthropic. "Claude Fable 5.1 & Claude Mythos 5.1 System Card," sections 8.4 and 8.5. September 1, 2026. https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf
15. OpenAI. "GPT-6 Astra: A new generation of intelligence." September 3, 2026. https://openai.com/index/gpt-6-astra/
16. OpenAI. "Introducing GPT-6 Sol and Luna." September 22, 2026. https://openai.com/index/introducing-gpt-6-sol-and-luna/
17. Cognition. "Introducing SWE-2: Pushing the Pareto Frontier." September 10, 2026. https://cognition.com/blog/swe-2
18. Google DeepMind. "Gemini 3.7 Flash - Model Card." Results as of August 2026. https://deepmind.google/models/model-cards/gemini-3-7-flash/
19. SpaceXAI. "Introducing Grok 4.6." August 12, 2026. https://x.ai/news/grok-4-6
20. Cognition. "Claude Opus 5.5 is now available in Devin." Devin blog, September 22, 2026. https://devin.ai/blog/claude-opus-5-5
21. Cognition. "Claude Sonnet 5.5 is now available in Devin." Devin blog, September 28, 2026. https://devin.ai/blog/claude-sonnet-5-5
22. Cognition (@cognition). "Introducing FrontierCode: a coding eval that raises the bar for difficulty & quality." X, June 8, 2026. https://x.com/cognition/status/2064061031912288715
23. Scott Wu (@ScottWu46). Post on FrontierCode and SWE-Bench-style grading. X, June 8, 2026. https://x.com/ScottWu46/status/2064073699368800475
24. swyx (@swyx). "It's finally out!!!" X, June 8, 2026. https://x.com/swyx/status/2064081945567580323
25. Theo (@theo). "This chart scares me a bit." and "Open questions for the @cognition team who worked on this." X, June 8, 2026. https://x.com/theo/status/2064107107356741953 and https://x.com/theo/status/2064126021088215385
26. Cognition (@cognition). Reply to @theo. X, June 9, 2026. https://x.com/cognition/status/2064215347503452649
27. Latent Space. "[AINews] FrontierCode: Benchmarking for Code Quality over Slop." June 9, 2026. https://www.latent.space/p/ainews-frontiercode-benchmarking
28. Parker Whitfill, Cheryl Wu, Joel Becker and Nate Rush. "Many SWE-bench-Passing PRs Would Not Be Merged into Main." METR, March 10, 2026. https://metr.org/notes/2026-03-10-many-swe-bench-passing-prs-would-not-be-merged-into-main

