CursorBench
CursorBench is an agentic coding benchmark built and run by Cursor, the AI code editor made by Anysphere. It scores coding agents on what Cursor calls "ambiguous, multi-file tasks from real Cursor sessions," and publishes the results on a public leaderboard that lists each model at each reasoning-effort setting together with its average cost, tokens and agent steps per task.[1] Cursor describes CursorBench as "our internal eval suite based on real Cursor sessions from our engineering team" and published its first dedicated write-up of it on March 11, 2026.[2] Cursor had first described the benchmark publicly, as "Cursor Bench," on October 29, 2025, in the launch post for its first Composer model, which reported results "on an internal benchmark in the Cursor tool harness."[33] The benchmark has gone through versions 3.0 (March 2026), 3.1 (May 2026), 3.2 (July 2026) and 4.0 (September 10, 2026), and Cursor says scores should only be compared within one version.[1][2]
Cursor uses CursorBench "throughout training and evaluation" of its own Composer models.[17] During 2026 it became a figure that other labs quote in their launch materials: Anthropic prints CursorBench rows in the benchmark tables for Claude Fable 5.1, Claude Opus 5.5 and Claude Sonnet 5.5, and SpaceXAI does the same for Grok 4.6 and Grok 4.7. Anthropic states that the scores were measured by Cursor, and the Grok figures match Cursor's leaderboard.[5][23][25][12][1] Cursor has been part of SpaceX since August 14, 2026. It trained Grok 4.5 jointly with SpaceXAI and describes Grok 4.6 as a model "we released," so the company that runs the benchmark also builds some of the models it ranks, including Grok models whose CursorBench scores appear in SpaceXAI's launch tables.[6][18] As of September 29, 2026, the CursorBench 4.0 leaderboard is led by Claude Opus 5.5 at max effort (57.8%), with Claude Sonnet 5.5 the second-highest model (55.5% at max effort, fourth entry overall behind three Opus 5.5 settings).[1]
Overview
| Property | Detail |
|---|---|
| Developer | Cursor (Anysphere, part of SpaceX since August 14, 2026)[6] |
| Type | Offline evaluation suite for coding agents[2] |
| First public description | October 29, 2025, as "Cursor Bench" in Cursor's Composer launch post;[33] first dedicated write-up March 11, 2026 ("How we compare model quality in Cursor" by Naman Jain)[2] |
| Current version | CursorBench 4.0, introduced September 10, 2026[1] |
| Task source | Real Cursor sessions, traced with Cursor Blame; many tasks from Cursor's internal codebase and "controlled sources"[2] |
| Grading | "Agentic graders" for solution correctness[2] |
| Execution | Cursor agents run in Cursor's production harness on Anyrun, Cursor's internal platform for sandboxed coding environments[3][17] |
| Reported metrics | Score, average cost per task, tokens per task, steps per task, at each effort level[1] |
| Public leaderboard | cursor.com/cursorbench[1] |
| Third-party tracking | Epoch AI mirrors the public leaderboard[7] |
Background
Cursor built CursorBench because it considered public coding benchmarks a poor guide to how models behave in its product. Its March 2026 post gives three reasons. The first is alignment: most SWE benchmarks focus on bug fixing, and Cursor says Terminal-Bench "emphasizes broad puzzle-style tasks like finding the best chess move from a board position." The second is grading: public tasks assume a narrow set of correct answers, while real developer requests are "underspecified enough to admit many valid approaches." The third is contamination: SWE-bench Verified, SWE-Bench Pro and SWE-bench Multilingual draw tasks from public repositories that end up in training data.[2]
The Composer 2 technical report, posted to arXiv in March 2026, sets out the same argument in four parts (domain mismatch, prompt over-specification, data contamination and overfitting, and narrow evaluation scope). As an example of compressed scores, it notes that Claude Haiku 4.5 scores 73.3% on SWE-bench Verified, close to GPT-5's 74.9%, which the report calls misaligned with accuracy on broader task distributions such as Terminal-Bench.[3] Cursor's blog makes a similar point, saying that on public benchmarks "models like Haiku can match or exceed GPT-5," while CursorBench "distinguishes reliably between models that developers experience as meaningfully different."[2]
Task design
Cursor sources tasks with Cursor Blame, a tool that "traces committed code back to the agent request that produced it," giving a pairing of a developer's query and a ground-truth solution. Cursor says "many tasks come from our internal codebase and controlled sources, which reduces the risk that models have seen them in training," and that it refreshes the suite "every few months."[2] Anthropic's system cards for Claude Sonnet 5 and Claude Fable 5.1 describe the tasks as "drawn from internal use and external traffic."[8][9]
The prompts are deliberately short. In the Composer 2 report, CursorBench tasks have a median of 181 lines changed in the reference diff, against 7 to 10 lines for SWE-bench Verified and Multilingual, and a median task description of 390 characters, against 1,185 to 3,055 characters for the public benchmarks.[3] The report says CursorBench-3 "more than doubles the median task size from the initial version," and that the mix has shifted as developers hand agents "long-running command execution, experiment monitoring, and data analysis."[3] Cursor's blog adds monorepo workspaces, production-log investigations and long-running experiments to the list of harder task types.[2]
The report gives two examples; the first is labelled as truncated and obfuscated. In one, the agent gets a terse bug report about a retry loop that reports failure after a later attempt succeeded, plus production logs that include unrelated warnings; the real cause is an esbuild 0.20.2 downleveling bug. In the other, the agent must write and tune a heuristic that counts malformed "growing prefix" streaming responses across 954 JSON log files.[3]
Scoring and harness
Cursor runs each task "exactly as it would execute in our production environment," using the same Anyrun infrastructure that hosts its reinforcement-learning environments. Accuracy is "aggregated over all tasks across multiple passes of the evaluation set to reduce variance," and Cursor also tracks completion tokens, latency and inference cost.[3] Cursor says its task descriptions are "intentionally short" and that it uses "agentic graders to reliably score them."[2]
The leaderboard reports a correctness score and three efficiency columns. Average cost per task "is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task." The page adds that "results are subject to variance; small differences in scores may not be statistically meaningful."[1] Under version 3.2, the cost note said costs were averaged "with the same task weights as the CursorBench 3.2 score," which indicates the score is a weighted aggregate.[11] Models appear at several effort levels (typically low to max, plus "Minimal" for Muse Spark 1.3), though not every model is listed at every level; Composer 2.5 appears as a single entry.[1]
Cursor treats CursorBench as the offline half of a "hybrid online-offline eval process." The online half consists of controlled experiments on live traffic, meant to catch cases where "the agent's output looks correct to a grader but feels worse to a developer using the product." Cursor says CursorBench rankings track its online metrics more closely than public benchmarks do.[2] The Composer 2 report lists further targeted evaluations run alongside it: intent, instruction following, eager editing, code quality and interruption handling.[3]
Versions
Cursor's leaderboard changelog lists four task versions and several reporting corrections. The Composer 2 report also refers to earlier iterations that preceded CursorBench-3.[1][3]
| Date | Type | Change (Cursor's wording) |
|---|---|---|
| March 11, 2026 | Tasks | CursorBench 3.0: "Initial set of tasks focused on edit, refactor, and bugfix problems." |
| May 19, 2026 | Tasks | CursorBench 3.1: "Introduced problems focused on codebase understanding, bugfinding, planning, and code review." Also "improved grading criteria for some edit tasks." |
| July 8, 2026 | Tasks | CursorBench 3.2: "Introduced instruction following and advanced tool use problems." |
| July 9, 2026 | Reporting | GPT-5.6 Sol, Terra and Luna results updated "to account for cache write costs." |
| July 30, 2026 | Reporting | GPT-5.6 Terra and Luna results updated "to account for adjusted pricing." |
| August 11, 2026 | Reporting | Sonnet 5 results updated "to account for adjusted pricing." |
| September 10, 2026 | Tasks | CursorBench 4.0: "new long-horizon problems focused on edit, refactor, investigation, intent understanding, managing jobs, and design adherence." |
Source: CursorBench changelog.[1]
Each version change resets the scale. An update to the March blog post says 3.1 scores "can shift from the numbers and charts in this post and should be compared within the same eval version."[2] When 4.0 arrived, Lee Robinson, whose X profile describes his role as "ML @SpaceXAI," posted "We just rolled out CursorBench 4.0!" and wrote that it "is more difficult than before (so all models score lower)."[13] Epoch AI archived its 3.x results "as 4.0 scores are not comparable with earlier versions," and Anthropic's system cards carry the note that "previous system cards reported older versions of CursorBench, so the scores are not comparable."[7][9][27]
The March blog post's own header still names 3.1 as "the current production version" as of September 29, 2026, while the leaderboard it links to shows 4.0.[2][1]
Cost columns can also change after publication. GPT-5.6 Sol at max effort was listed at $5.22 per task on a copy of the 3.2 leaderboard archived on July 9, 2026 and at $5.69 on a copy archived on September 9; Sonnet 5 at max effort went from $6.45 to $4.30 over the same period. The changelog lists a July 9 update for cache-write costs and an August 11 pricing adjustment for Sonnet 5. The scores themselves did not change.[11][12][1]
Results
The tables below list results from each version separately. Scores from different versions are on different task sets and should not be compared or ranked together.
CursorBench 3 (Composer 2 technical report, March 2026)
The Composer 2 report is the earliest source with a multi-model table. It identifies the results as CursorBench-3 and reports them without effort labels for most models.[3]
| Model | CursorBench-3 |
|---|---|
| GPT-5.4 | 63.9 |
| Composer 2 | 61.3 |
| GPT-5.3 Codex | 59.1 |
| Claude Opus 4.6 (high) | 58.2 |
| GPT-5.2 | 56.5 |
| Claude Opus 4.5 (high) | 48.4 |
| Composer 1.5 | 44.2 |
| GLM-5 | 42.7 |
| Composer 1 | 38.0 |
| Kimi K2.5 | 36.0 |
Source: Composer 2 Technical Report, Table 1.[3]
CursorBench 3.1 (leaderboard, June 2026)
Selected rows from a copy of the leaderboard archived on June 29, 2026, the last 3.1 copy before version 3.2 replaced it.[14]
| Model and effort | Score | Cost per task |
|---|---|---|
| Claude Fable 5, max | 72.9% | $18.02 |
| Claude Opus 4.7, max | 64.8% | $11.02 |
| GPT-5.5, extra high | 64.3% | $4.37 |
| Claude Opus 4.8, max | 63.8% | $7.59 |
| Composer 2.5 | 63.2% | $0.55 |
| GLM 5.2, max | 54.6% | $3.11 |
| Composer 2 | 52.2% | $0.56 |
| Gemini 3.5 Flash | 49.8% | $1.94 |
| Claude Sonnet 4.6, max | 49.0% | $3.09 |
| Kimi 2.6 | 47.6% | $1.27 |
| Kimi 2.5 | 31.9% | $0.87 |
Source: archived CursorBench leaderboard, June 29, 2026.[14]
CursorBench 3.2 (leaderboard, July to September 2026)
Selected rows, showing each model's best setting, from a copy archived on September 9, 2026, the day before 4.0 replaced it.[12]
| Model and effort | Score | Cost per task |
|---|---|---|
| Claude Fable 5.1, max | 73.4% | $9.64 |
| Grok 4.6, extra high | 70.8% | $2.81 |
| Claude Fable 5, max | 70.5% | $17.32 |
| Claude Opus 5, max | 70.0% | $8.23 |
| Gemini 3.8 Flash, high | 69.2% | $2.38 |
| Muse Spark 1.3, max | 67.9% | $1.31 |
| GPT-5.6 Sol, max | 67.2% | $5.69 |
| GPT-5.6 Terra, max | 64.9% | $2.31 |
| Claude Opus 4.8, max | 62.3% | $5.77 |
| Claude Sonnet 5, max | 61.5% | $4.30 |
| GPT-5.6 Luna, max | 61.1% | $0.39 |
| Kimi K3, max | 60.8% | $2.70 |
| GPT-5.5, high | 58.4% | $2.05 |
| Composer 2.5 | 56.1% | $0.44 |
| GLM 5.2, max | 55.0% | $1.76 |
Source: archived CursorBench leaderboard, September 9, 2026.[12]
Earlier copies of the 3.2 leaderboard also listed Grok 4.5 at 66.7% (high), 65.4% (medium) and 63.5% (low), each marked with an asterisk for the contamination problem described below.[11] Those rows are absent from archived copies from August 13, 2026 onward.[15]
CursorBench 4.0 (current leaderboard)
Best setting for each model on the live leaderboard as of September 29, 2026.[1]
| Rank of best entry | Model and effort | Score | Cost per task | Tokens per task | Steps per task |
|---|---|---|---|---|---|
| 1 | Claude Opus 5.5, max | 57.8% | $13.43 | 218,363 | 185 |
| 4 | Claude Sonnet 5.5, max | 55.5% | $9.67 | 271,920 | 170 |
| 7 | Claude Fable 5.1, max | 51.8% | $17.28 | 117,236 | 128 |
| 12 | Claude Opus 5, max | 46.6% | $11.95 | 85,384 | 106 |
| 13 | Grok 4.7, extra high | 46.3% | $6.01 | 70,141 | 88 |
| 20 | GLM 5.3, max | 42.6% | $5.05 | 96,387 | 166 |
| 21 | GPT-5.6 Sol, max | 41.7% | $8.23 | 42,944 | 99 |
| 23 | Muse Spark 1.3, max | 41.6% | $2.64 | 52,005 | 98 |
| 24 | Grok 4.6, extra high | 41.4% | $6.10 | 49,814 | 56 |
| 25 | GPT-5.6 Terra, max | 41.3% | $5.14 | 60,814 | 107 |
| 28 | Gemini 3.8 Flash, high | 39.6% | $4.70 | 162,565 | 324 |
| 34 | GLM 5.3 Flash, max | 36.8% | $0.39 | 56,410 | 118 |
| 36 | GPT-5.6 Luna, max | 35.9% | $1.03 | 87,284 | 208 |
| 39 | Claude Sonnet 5, max | 34.1% | $7.17 | 149,257 | 140 |
| 55 | Composer 2.5 | 27.7% | $0.68 | 17,347 | 41 |
Source: CursorBench leaderboard, cursor.com/cursorbench, accessed September 29, 2026.[1]
The two newest Anthropic models by effort setting:
| Effort | Claude Opus 5.5 | Claude Sonnet 5.5 |
|---|---|---|
| Low | 43.7% ($1.17) | 35.8% ($0.50) |
| Medium | 52.5% ($2.91) | 39.2% ($0.70) |
| High | 56.0% ($3.97) | 47.8% ($1.67) |
| Extra high | 56.0% ($6.98) | 53.1% ($3.88) |
| Max | 57.8% ($13.43) | 55.5% ($9.67) |
Source: CursorBench leaderboard.[1] Anthropic's launch charts show the same scores; the per-task costs for the model being launched are Anthropic's estimates from token counts Cursor supplied.[5][27]
Use in model launches
CursorBench figures reach the public in three ways: Cursor's own model posts, Cursor executives' quotes in other labs' launch posts, and benchmark tables that labs build from Cursor's measurements.
| Date | Announcement | Version | CursorBench content |
|---|---|---|---|
| March 19, 2026 | Cursor, Composer 2 | Not named in post (the technical report identifies CursorBench-3) | Composer 2 61.3, Composer 1.5 44.2, Composer 1 38.0[16][17][3] |
| April 16, 2026 | Anthropic, Claude Opus 4.7 | Not named | Quote from Cursor CEO Michael Truell: Opus 4.7 is "clearing 70% versus Opus 4.6 at 58%"[19] |
| May 18, 2026 | Cursor, Composer 2.5 | 3.1 | Composer 2.5 at 63.2%, reported by The Decoder and The New Stack from Cursor's charts[28][29][30] |
| May 28, 2026 | Anthropic, Claude Opus 4.8 | Not named | Truell: Opus 4.8 "exceeds prior Opus models across every effort level"[20] |
| June 9, 2026 | Anthropic, Claude Fable 5 | Not named | Truell: "Claude Fable 5 is the state of the art model on CursorBench"[21] |
| June 30, 2026 | Anthropic, Claude Sonnet 5 system card | Not named | Sonnet 5 61.2%, Sonnet 4.6 49%, Opus 4.8 63.8%[8] |
| July 8, 2026 | Cursor, Grok 4.5 | 3.2 | CursorBench left out of the launch comparison because of training-data contamination[18] |
| July 24, 2026 | Anthropic, Claude Opus 5 | 3.2 | Opus 5 at max effort "within 0.5% of Fable 5's peak score, but at half the cost per task"[22] |
| August 12, 2026 | SpaceXAI, Grok 4.6 | 3.2 | Grok 4.6 (high) 69.9%, Grok 4.5 (high) 66.7%, GPT-5.6 Sol (max) 67.2%, Claude Fable 5 (max) 70.5%[23] |
| September 1, 2026 | Anthropic, Claude Fable 5.1 | 3.2.0 | Fable 5.1 73.4%, Fable 5 70.5%, Opus 5 70.0%, GPT-5.6 Sol 67.2%[24][9] |
| September 21, 2026 | SpaceXAI, Grok 4.7 | 4.0 | Grok 4.7 (xhigh) 46.3%, Grok 4.6 (high) 40.4%, GPT-5.6 Sol (max) 41.7%, Claude Fable 5.1 (max) 51.8%[25] |
| September 22, 2026 | Anthropic, Claude Opus 5.5 | 4.0 | Opus 5.5 57.8%, Fable 5.1 51.8%, Opus 5 46.6%, GPT-5.6 Sol 41.7%[26] |
| September 28, 2026 | Anthropic, Claude Sonnet 5.5 | 4.0 | Sonnet 5.5 55.5%, Sonnet 5 34.1%, Opus 5.5 57.8%[4] |
The Sonnet 5 system card's figures for Opus 4.8 (63.8%) and Sonnet 4.6 (49%) match the max-effort entries on Cursor's 3.1 leaderboard at the time.[8][14] The Opus 4.7 quote came while 3.0 was the current version; the 58% figure for Opus 4.6 is close to the 58.2 that the Composer 2 report gives Opus 4.6 (high) on CursorBench-3.[19][3]
How the labs present it
Anthropic's recent system cards say CursorBench is executed "end-to-end in Cursor's production agent harness"[5][9] and that the "scores were measured and reported independently by Cursor."[5][27] For the model being launched, the scores come from results Cursor sent to Anthropic directly: the Opus 5.5 card says they are "from a result file Cursor sent directly to us," and the Sonnet 5.5 card says "a results table Cursor sent directly to us." In both cases Anthropic computed that model's cost per task itself from Cursor's token counts at list prices, while other models' scores and costs are "as Cursor published them on its public leaderboard in September 2026."[27][5] The cards also note that "neither source includes uncertainty intervals" and that the chart's y-axis "begins at 23%, not at zero, so differences between points are visually magnified."[27][5]
Anthropic's launch pages lean on the cost dimension. The Opus 5.5 page says that at default (medium) effort Opus 5.5 scores 52.5%, "compared to 51.8% for Fable 5.1 (max) and 46.6% for Opus 5 (max)," and "beats GPT-5.6 Sol's top score (41.7%) by 11 points for about a third of the cost per task."[26] The Sonnet 5.5 page says Sonnet 5.5 at low effort "exceeds Sonnet 5's best score for less than a tenth of the cost per task" and that its best score "is within about two points of Opus 5.5." Because GPT-6 Sol was not on the leaderboard, Anthropic wrote that "CursorBench 4.0 does not report GPT-6 Sol performance publicly, so we report GPT-5.6 Sol here."[4]
SpaceXAI's Grok 4.7 post says that on CursorBench 4.0, "which stresses longer-running coding tasks, Grok 4.7 is at the frontier in price-performance," and its headline table sets Grok 4.7 at extra-high effort against competitors at max.[25] Its Grok 4.6 table compared Grok 4.6 at high effort against competitors at max; Cursor's own 3.2 leaderboard listed Grok 4.6 higher at extra-high effort (70.8%).[23][12]
Several launch quotes come from Cursor's leadership. Michael Truell, Cursor's co-founder and CEO, supplied the CursorBench quotes for Opus 4.7, Opus 4.8 and Fable 5.[19][20][21] Sualeh Asif is quoted as "Co-Founder" in Anthropic's Opus 5 post, in a quote about Cursor, and as "Director of ML" at SpaceXAI in the Fable 5.1 and Sonnet 5.5 posts. His Sonnet 5.5 quote reads: "Claude Sonnet 5.5 delivers frontier-level performance on CursorBench 4.0 at 55.5%, second only to Opus 5.5."[22][24][4]
Contamination and criticism
Grok 4.5 contamination
The Composer 2 report said that because CursorBench tasks come from real agent sessions, the benchmark reflects real work "while completely avoiding train-set contamination."[3] That claim did not hold for Grok 4.5, which Cursor trained jointly with SpaceXAI on "trillions of tokens of Cursor data." Cursor's Grok 4.5 launch post, published July 8, 2026, says: "Grok 4.5 has an advantage on CursorBench because an earlier snapshot of the Cursor codebase was accidentally included in training. The exact impact is unclear. That data has been removed for future models, and in parallel we are working on a larger update to CursorBench, hence the exclusion here."[18] The leaderboard carried the three Grok 4.5 rows with an asterisk and a similar note, and the rows had been removed by August 13, 2026.[11][15]
The exclusion did not reach every later comparison. SpaceXAI's Grok 4.6 launch table of August 12, 2026 lists Grok 4.5 (high) at 66.7% on CursorBench v3.2, the asterisked figure, as the baseline for Grok 4.6's 69.9%, and the page carries no contamination note.[23][11]
Independence
Cursor is both the operator of CursorBench and a model developer. Its first-party Composer models appear on the leaderboard next to models from Anthropic, OpenAI, Google and others, and since April 2026 it has partnered with SpaceX on model training, a process that ended with SpaceX completing its acquisition of Cursor on August 14, 2026.[32][6][1] The results have not been reproduced outside Cursor: Epoch AI's CursorBench page says it sources results "from the public CursorBench leaderboard."[7]
Composer's position has moved a lot between versions. Composer 2.5 ranked third of 14 entries on the 3.1 leaderboard in May 2026 (63.2%), was 27th of 42 on 3.2 in July (56.1%), and is 55th of 63 on 4.0 (27.7%).[10][11][1] When Composer 2.5 launched, The New Stack described the 63.2% as coming from Cursor's "own CursorBench v3.1" and wrote that benchmarks "don't provide any real assurance for how these models will perform in the real world."[30]
Cursor's case for private tasks
Cursor argues that its task sourcing is the answer to a different contamination problem. In a June 2026 study, it reported that 63% of successful Claude Opus 4.8 (max) resolutions on SWE-bench Pro "retrieved the fix rather than derived it," mostly by finding the merged pull request on the public web or mining bundled git history. When it sealed git history and restricted internet access, Opus 4.8 fell from 87.1% to 73.0% and Composer 2.5 from 74.7% to 54.0%. Cursor wrote that this is "one reason we prefer evals built from non-public repositories, like CursorBench," which "can test agentic coding ability while still letting agents use tools in the ways they would during real work."[31] The same post notes that Cursor's own Composer 2.5 had the largest SWE-bench Pro gap in the study, and that Cursor does not treat its standard SWE-bench Pro score "as a reliable benchmark number for Composer."[31] See reward hacking and data contamination.
Other limitations
- Cursor's leaderboard states that "small differences in scores may not be statistically meaningful," and Anthropic notes that Cursor's published results carry no uncertainty intervals.[1][27]
- The task set is internal, and the leaderboard and March blog post do not give a task count.[1][2]
- Cursor's blog lists code quality, efficiency and interaction behavior as dimensions it evaluates, but the public leaderboard reports only correctness plus cost, tokens and steps.[2][1]
- Several 2026 launch quotes give a CursorBench score without naming the version, even though the scale resets with each version.[1][19][21]
See also
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22Cursor. "CursorBench." cursor.com. Accessed September 29, 2026. cursor.com/cursorbench
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16Jain, Naman. "How we compare model quality in Cursor." Cursor blog, March 11, 2026 (updated May 2026). cursor.com/...cursorbench
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13Cursor Research. "Composer 2 Technical Report." arXiv:2603.24477, March 2026. arxiv.org/...2603.24477
- ^1 ^2 ^3Anthropic. "Introducing Claude Sonnet 5.5." anthropic.com, September 28, 2026. anthropic.com/claude-sonnet-5-5
- ^1 ^2 ^3 ^4 ^5 ^6Anthropic. "Claude Sonnet 5.5 System Card." September 28, 2026. Section 8.8. www-cdn.anthropic.com/...205.5%20System%20Card.pdf
- ^1 ^2 ^3Cursor Team. "Cursor is now a part of SpaceX." Cursor blog, August 14, 2026. cursor.com/...joining-spacex
- ^1 ^2 ^3Epoch AI. "CursorBench 4.0." epoch.ai. Accessed September 29, 2026. epoch.ai/...cursorbench
- ^1 ^2 ^3Anthropic. "Claude Sonnet 5 System Card." June 30, 2026. Section 8.5. www-cdn.anthropic.com/...t%205%20System%20Card.pdf
- ^1 ^2 ^3 ^4Anthropic. "Claude Fable 5.1 & Claude Mythos 5.1 System Card." September 1, 2026. Section 8.8. www-cdn.anthropic.com/...205.1%20System%20Card.pdf
- ^Cursor. "CursorBench" (CursorBench 3.1 leaderboard). Internet Archive copy of May 20, 2026. web.archive.org/...cursorbench
- ^1 ^2 ^3 ^4 ^5 ^6Cursor. "CursorBench" (CursorBench 3.2 leaderboard). Internet Archive copy of July 9, 2026. web.archive.org/...cursorbench
- ^1 ^2 ^3 ^4 ^5Cursor. "CursorBench" (CursorBench 3.2 leaderboard). Internet Archive copy of September 9, 2026. web.archive.org/...cursorbench
- ^Robinson, Lee (@leerob). Post on X announcing CursorBench 4.0, September 10, 2026. x.com/...2098144600594465148
- ^1 ^2 ^3Cursor. "CursorBench" (CursorBench 3.1 leaderboard). Internet Archive copy of June 29, 2026. web.archive.org/...cursorbench
- ^1 ^2Cursor. "CursorBench" (CursorBench 3.2 leaderboard). Internet Archive copy of August 13, 2026. web.archive.org/...cursorbench
- ^Cursor Team. "Introducing Composer 2." Cursor blog, March 19, 2026. cursor.com/...composer-2
- ^1 ^2 ^3Rush, Sasha. "A technical report on Composer 2." Cursor blog, March 27, 2026. cursor.com/...composer-2-technical-report
- ^1 ^2 ^3Cursor Team. "Introducing Grok 4.5." Cursor blog, July 8, 2026. cursor.com/...grok-4-5
- ^1 ^2 ^3 ^4Anthropic. "Introducing Claude Opus 4.7." anthropic.com, April 16, 2026. anthropic.com/...claude-opus-4-7
- ^1 ^2Anthropic. "Introducing Claude Opus 4.8." anthropic.com, May 28, 2026. anthropic.com/...claude-opus-4-8
- ^1 ^2 ^3Anthropic. "Claude Fable 5 and Claude Mythos 5." anthropic.com, June 9, 2026. anthropic.com/...claude-fable-5-mythos-5
- ^1 ^2Anthropic. "Introducing Claude Opus 5." anthropic.com, July 24, 2026. anthropic.com/...claude-opus-5
- ^1 ^2 ^3 ^4SpaceXAI. "Introducing Grok 4.6." x.ai, August 12, 2026. x.ai/...grok-4-6
- ^1 ^2Anthropic. "Introducing Claude Fable 5.1 and Claude Mythos 5.1." anthropic.com, September 2026. anthropic.com/claude-fable-and-mythos-5-1
- ^1 ^2 ^3SpaceXAI. "Introducing Grok 4.7." x.ai, September 21, 2026. x.ai/...grok-4-7
- ^1 ^2Anthropic. "Introducing Claude Opus 5.5." anthropic.com, September 22, 2026. anthropic.com/claude-opus-5-5
- ^1 ^2 ^3 ^4 ^5 ^6Anthropic. "Claude Opus 5.5 System Card." September 22, 2026. Section 8.8. www-cdn.anthropic.com/...205.5%20System%20Card.pdf
- ^Cursor Team. "Introducing Composer 2.5." Cursor blog, May 18, 2026. cursor.com/...composer-2-5
- ^Bastian, Matthias. "Cursor's Composer 2.5 matches Opus 4.7 and GPT-5.5 benchmarks at a fraction of the cost." The Decoder, May 18, 2026. the-decoder.com/...marks-at-a-fraction-of-the-cost
- ^1 ^2Shubel, Meredith. "Cursor bets on cheaper coding with Composer 2.5 and Kimi K2.5." The New Stack, May 20, 2026. thenewstack.io/cursor-composer-benchmarks
- ^1 ^2Jain, Naman. "Reward hacking is swamping model intelligence gains." Cursor blog, June 25, 2026. cursor.com/...reward-hacking-coding-benchmarks
- ^Cursor Team. "Cursor partners with SpaceX on model training." Cursor blog, April 21, 2026. cursor.com/...spacex-model-training
- ^1 ^2Cursor Team. "Composer: Building a fast frontier model with RL." Cursor blog, October 29, 2025. cursor.com/...composer
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,467 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: xg12 V3 independent verification 29 Sep 2026: ~62 sources, all 4.0/3.2/3.1 table cells checked vs leaderboard and Wayback; 3 material (first-public date, Grok 4.6 co-release) + 5 minor fixed.
Cite this page: AI Wiki. "CursorBench." aiwiki.ai, updated 29 Sept 2026, fact-checked 29 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/cursorbench