# Best AI Coding Assistants

> Source: https://aiwiki.ai/wiki/best_ai_coding_assistants
> Updated: 2026-09-23
> Fact-checked: 2026-09-23
> Categories: AI Agents, AI Code Generation, Developer Tools, Large Language Models
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Best AI Coding Assistants." aiwiki.ai, 23 Sept 2026. https://aiwiki.ai/wiki/best_ai_coding_assistants
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

As of 23 September 2026, the two tools at the top of the official [Terminal-Bench](https://aiwiki.ai/wiki/terminal_bench) 4.0 leaderboard are OpenAI's Codex and Anthropic's [Claude Code](https://aiwiki.ai/wiki/claude_code). Codex running [GPT-6 Astra](https://aiwiki.ai/wiki/gpt_6_astra) at max effort scores 58.2 percent, and Claude Code running [Claude Fable 5.1](https://aiwiki.ai/wiki/claude_fable_5_1) at max effort scores 57.9 percent, a gap well inside the board's roughly 3-point confidence intervals [31]. Neither Anthropic's [Claude Opus 5.5](https://aiwiki.ai/wiki/claude_opus_5_5), released on 22 September 2026 and now the default model in Claude Code, nor OpenAI's [GPT-6 Sol](https://aiwiki.ai/wiki/gpt_6_sol), released the same day, had a public leaderboard entry yet; Anthropic reports 66.4 percent for Opus 5.5 on its own reproduction of the benchmark, and the independent evaluator [Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis) measures Opus 5.5 and GPT-6 Astra level at 59.6 percent with a single shared harness [35][41]. For in-editor (IDE) coding, [Cursor](https://aiwiki.ai/wiki/cursor), now owned by SpaceX, and [GitHub Copilot](https://aiwiki.ai/wiki/github_copilot) remain the main choices, and Cognition has folded [Windsurf](https://aiwiki.ai/wiki/windsurf) into its Devin Desktop editor alongside the [Devin](https://aiwiki.ai/wiki/devin) cloud agent [45][50].

## Quick verdict: which AI coding assistant is best?

- Highest ceiling on hard, autonomous, multi-file work: Claude Code (Opus 5.5 by default, Claude Fable 5.1 on request) or OpenAI Codex (GPT-6 Astra). They hold the top five places on Terminal-Bench 4.0 between them [31].
- Best value at the top of the board: [OpenAI Codex](https://aiwiki.ai/wiki/openai_codex) with GPT-6 Astra, whose max-effort leaderboard run cost about $3.3k against about $6.2k for Claude Code with Fable 5.1 at max effort, at essentially the same score [31]. Codex is included in every ChatGPT plan, from Free upward [10].
- Best AI IDE (in-editor flow): Cursor, which pairs its first-party Grok and Composer models with frontier models from Anthropic, OpenAI, and Google [44].
- Best for the widest team and enterprise rollout: GitHub Copilot, which runs in VS Code, Visual Studio, JetBrains IDEs, Xcode, Neovim, Eclipse, Zed, and other editors and offers models from OpenAI, Anthropic, Google, and xAI [16][18].
- Best fully autonomous cloud engineer: Devin, with Cognition's new SWE-2 model [51].
- Best open-source and model-agnostic: [Aider](https://aiwiki.ai/wiki/aider) (terminal) and [Cline](https://aiwiki.ai/wiki/cline) (IDE and CLI) [25][26].
- Free terminal option: [Gemini CLI](https://aiwiki.ai/wiki/gemini_cli) no longer is one. Google ended free personal-account access on 18 June 2026 and points individual users to Antigravity CLI [42][43].

## Summary comparison table

| Tool | Developer | Default / notable models | SWE-bench Verified (vals.ai) | Terminal-Bench 4.0 (tbench.ai) | Price (USD) | Type / access |
|------|-----------|--------------------------|------------------------------|--------------------------------|-------------|---------------|
| [Claude Code](https://aiwiki.ai/wiki/claude_code) | [Anthropic](https://aiwiki.ai/wiki/anthropic) | [Opus 5.5](https://aiwiki.ai/wiki/claude_opus_5_5) default; [Fable 5.1](https://aiwiki.ai/wiki/claude_fable_5_1) selectable | 97.0 (Opus 5) / 95.0 (Fable 5) [3] | 57.9 (Fable 5.1) / 53.9 (Opus 5) [31] | Pro $20, Max $100 or $200/mo, or API [5][62] | Terminal agent + IDE, desktop, web; proprietary [54][55] |
| [OpenAI Codex CLI](https://aiwiki.ai/wiki/codex_cli) | [OpenAI](https://aiwiki.ai/wiki/openai) | [GPT-6 Sol](https://aiwiki.ai/wiki/gpt_6_sol) recommended; [GPT-6 Astra](https://aiwiki.ai/wiki/gpt_6_astra) | 96.2 (GPT-5.6 Sol) [3] | 58.2 (GPT-6 Astra) / 37.3 (GPT-5.6 Sol) [31] | Free; Go $8, Plus $20, Pro from $100, Business $20/user; or API [10] | Terminal agent + IDE + desktop app + cloud; CLI open source ([Apache-2.0](https://aiwiki.ai/wiki/apache_license)) [13] |
| [Cursor](https://aiwiki.ai/wiki/cursor) | [Anysphere](https://aiwiki.ai/wiki/anysphere) (SpaceX) | [Grok 4.7](https://aiwiki.ai/wiki/grok_4_7), [Composer 2.5](https://aiwiki.ai/wiki/composer_2_5) + frontier | 79.6 (Composer 2.5) [3] | n/r | Hobby free; Pro $20, Pro+ $60, Ultra $200/mo; Teams $40/user [14] | AI IDE (VS Code codebase), CLI, cloud agents [48] |
| [GitHub Copilot](https://aiwiki.ai/wiki/github_copilot) | GitHub | Multi-model (Auto; GPT-6, Claude, Gemini, Grok) | up to 97.0 (via Opus 5) [3] | n/r | Free; Pro $10, Pro+ $39, Max $100/mo; Business $19, Enterprise $39/user [16] | IDE ext + agent + CLI |
| Devin Desktop (formerly [Windsurf](https://aiwiki.ai/wiki/windsurf)) | [Cognition](https://aiwiki.ai/wiki/cognition_ai) | SWE-2 + frontier | n/r | n/r | Free; Pro $20, Max $200/mo; Teams $80 + $40/seat [20] | AI IDE with agent manager |
| [Devin](https://aiwiki.ai/wiki/devin) | [Cognition](https://aiwiki.ai/wiki/cognition_ai) | SWE-2 + frontier | n/r | n/r | Cloud agents from Pro $20/mo [20] | Autonomous cloud agent |
| [Grok Build](https://aiwiki.ai/wiki/grok_build) | [SpaceXAI](https://aiwiki.ai/wiki/xai) | [Grok 4.7](https://aiwiki.ai/wiki/grok_4_7) | n/r | 37.6 (Grok 4.7) [31] | Free to try; Grok API [46] | Terminal agent |
| [Gemini CLI](https://aiwiki.ai/wiki/gemini_cli) | Google | Gemini 3 models | 78.8 (Gemini 3.1 Pro) [3] | n/r | Paid Gemini API key or Code Assist Standard/Enterprise [42] | Terminal agent, open source ([Apache-2.0](https://aiwiki.ai/wiki/apache_license)) [22] |
| [Aider](https://aiwiki.ai/wiki/aider) | Aider-AI (open source) | Bring your own model | inherits model | n/r | Free, pay model API | Terminal pair programmer, open source ([Apache-2.0](https://aiwiki.ai/wiki/apache_license)) [25] |
| [Cline](https://aiwiki.ai/wiki/cline) | Cline Bot Inc. (open source) | Bring your own model | inherits model | n/r | Free, pay model API | SDK, IDE ext, CLI, open source ([Apache-2.0](https://aiwiki.ai/wiki/apache_license)) [26] |
| [Augment Code](https://aiwiki.ai/wiki/augment_code) | Augment | Context Engine + LLMs (Opus 4.5 in its SWE-bench Pro run) | n/r (self-reports 65.4, March 2025 [28]) | n/r | Standard $20/mo, Business $100/mo (flat, up to 50 seats), Enterprise custom [27] | CLI + cloud agents |

Last verified: 23 September 2026. Prices in USD. The SWE-bench Verified column uses vals.ai's independent results with a single bash-only harness; vals.ai has archived that benchmark as saturated (its results were last updated on 1 September 2026) and does not test newer models such as Opus 5.5, Fable 5.1, GPT-6 Astra, or GPT-6 Sol [3]. The Terminal-Bench 4.0 column is a tool-plus-model score from the official leaderboard, best result across reasoning-effort settings [31]. An n/r cell means no current primary source reports that figure. Bring-your-own-model tools inherit the score of whichever model you attach.

## Best terminal coding agents (CLI)

Command-line agents run in your terminal, read and edit the whole repository, run tests, and iterate on failures. They hold every place on the Terminal-Bench 4.0 leaderboard: the board lists entries for Codex, Claude Code, Grok Build, and the minimal mini-SWE-agent harness [31].

### 1. Claude Code (Anthropic)

Best for: complex, autonomous, multi-file engineering where correctness matters most.

[Claude Code](https://aiwiki.ai/wiki/claude_code) is [Anthropic](https://aiwiki.ai/wiki/anthropic)'s agentic coding tool. It runs in the terminal, in VS Code and JetBrains IDEs, in a desktop app, and on the web [54]. It is proprietary software: its GitHub repository [52] carries an "All rights reserved" notice [55]. Since version 2.1.280 the default model on Pro, Max, Team, Enterprise, and the Anthropic API is Claude Opus 5.5; before that, Pro and Team Standard defaulted to Claude Sonnet 5 and the other plans to [Claude Opus 5](https://aiwiki.ai/wiki/claude_opus_5) [4]. Claude Fable 5.1 and Fable 5 are selectable but are not the default on any plan [4]. Fable 5 was briefly withdrawn after US export controls were applied on 12 June 2026, returned on 1 July, and moved to usage credits on subscription plans after 7 July [7]. On claude.com's plan table, Fable runs on usage credits on Pro and within 50 percent of weekly limits on Max [5].

On the public Terminal-Bench 4.0 board, Claude Code scores 57.9 percent with Fable 5.1 and 53.9 percent with Opus 5 [31]. Opus 5.5 was not yet listed on 23 September; Anthropic reports 66.4 percent for it at xhigh effort on a setup that reproduces the leaderboard's Opus 5 result within noise, and Artificial Analysis measures it at 59.6 percent [35][41]. Anthropic says Opus 5.5 performs at the level of Fable 5.1 on most work and costs 40 percent less to run than Opus 5 [35]. Plans: Pro $20 per month ($17 with annual billing), Max $100 per month (5x Pro usage) or $200 per month (20x), plus API pay-as-you-go [5][62]. API token prices per million are $4 input and $20 output for Opus 5.5, $10 and $50 for Fable 5.1, and $5 and $25 for Opus 5, which launched on 24 July 2026 [5][37]. Fable 5.1 kept Fable 5's $10 and $50 pricing [36]. With Opus 5.5, Anthropic also raised five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans [35].

### 2. OpenAI Codex CLI (OpenAI)

Best for: top-of-board agentic coding at lower cost, and anyone already paying for ChatGPT.

The [OpenAI Codex CLI](https://aiwiki.ai/wiki/codex_cli) is [OpenAI](https://aiwiki.ai/wiki/openai)'s open-source terminal agent, released under the [Apache-2.0](https://aiwiki.ai/wiki/apache_license) license, and part of a Codex product that also runs in the ChatGPT desktop app, IDE extensions, and the cloud [13][9]. OpenAI released GPT-6 Astra on 3 September 2026 and GPT-6 Sol and GPT-6 Luna on 22 September 2026 [38][39][40]. Codex's documentation recommends GPT-6 Sol for complex coding and agentic work and GPT-6 Astra for the hardest end-to-end tasks; Astra is available in the Codex CLI and IDE extension but not in Codex cloud [9]. GPT-5.5 retires from Codex for ChatGPT sign-ins on 14 October 2026 [9].

Codex with GPT-6 Astra holds first place on Terminal-Bench 4.0 at 58.2 percent (max effort), and ties Claude Code with Fable 5.1 at 57.9 percent at high and xhigh effort [31]. The previous-generation GPT-5.6 Sol, released on 9 July 2026, scores 37.3 percent in Codex [31][40]. OpenAI's own GPT-5.5 launch in April reported 82.7 percent on the older Terminal-Bench 2.0 and 58.6 percent on [SWE-bench Pro](https://aiwiki.ai/wiki/swe_bench_pro), without a SWE-bench Verified figure; vals.ai later measured GPT-5.5 at 82.6 percent on Verified [11][3]. Codex is included in ChatGPT Free, Go ($8), Plus ($20), Pro (from $100 per month), and Business ($20 per user per month billed annually) [10]. API prices per million tokens are $10 input and $50 output for GPT-6 Astra, $2 and $10 for GPT-6 Sol, and $5 and $30 for [GPT-5.5](https://aiwiki.ai/wiki/gpt-5.5) [12][38][39]. OpenAI cut Sol and Luna prices by half compared with GPT-5.6's promotional pricing [39].

### 3. Grok Build (SpaceXAI)

Best for: developers already using Grok models, and a free way to try them in the terminal.

[Grok Build](https://aiwiki.ai/wiki/grok_build) is SpaceXAI's coding agent, installed from x.ai/build [46]. With [Grok 4.7](https://aiwiki.ai/wiki/grok_4_7), released on 21 September 2026, it scores 37.6 percent on Terminal-Bench 4.0 (over 330 trials; the leaderboard marks only its cost figure as partial, covering 324 of 330 trials), up from 20.3 percent with Grok 4.6 [31][46]. SpaceXAI's launch page first gave Grok 4.7 38.0 percent on Terminal-Bench 4.0 and later changed the figure to 37.6 percent, which matches the leaderboard [46]. Grok 4.7 costs $2 per million input tokens and $6 per million output tokens on the Grok API, and SpaceXAI invites users to try it in Grok Build for free [46].

### 4. Gemini CLI and Antigravity CLI (Google)

Best for: teams that already hold Gemini Code Assist Standard or Enterprise licenses or paid Gemini API keys.

[Gemini CLI](https://aiwiki.ai/wiki/gemini_cli) is Google's open-source ([Apache-2.0](https://aiwiki.ai/wiki/apache_license)) terminal agent [22]. Its main draw used to be a free tier through a personal Google account (60 requests per minute and 1,000 per day) [22], but on 19 May 2026 Google announced that it was unifying its terminal tools into [Antigravity](https://aiwiki.ai/wiki/antigravity) and its new Antigravity CLI [42]. From 18 June 2026, Gemini CLI stopped serving requests for Gemini Code Assist for individuals, Google AI Pro, and Google AI Ultra, and the "Login with Google" option no longer works; affected users are directed to Antigravity and Antigravity CLI [43][30]. Gemini CLI remains available through paid Gemini API keys and through Code Assist Standard and Enterprise licenses [42]. Google's model card gives [Gemini 3.1 Pro](https://aiwiki.ai/wiki/gemini_3_1_pro) 80.6 percent on SWE-bench Verified, while vals.ai's harness measures the preview at 78.8 percent [8][3]. On the last Terminal-Bench 2.1 board, Gemini CLI with Gemini 3.1 Pro scored 65.8 percent; it has no Terminal-Bench 4.0 entry, and Google's [Gemini 3.8 Flash](https://aiwiki.ai/wiki/gemini_3_8_flash) appears there only with the mini-SWE-agent harness, at 19.1 percent [1][31]. Gemini API prices are $2 input and $12 output per million tokens for Gemini 3.1 Pro Preview (prompts up to 200K tokens) and $0.75 and $3.75 for Gemini 3.8 Flash through 31 December 2026 [23].

### 5. Aider (open source)

Best for: a lightweight, scriptable, model-agnostic pair programmer in the terminal.

[Aider](https://aiwiki.ai/wiki/aider) is a free, open-source ([Apache-2.0](https://aiwiki.ai/wiki/apache_license)) terminal tool that edits your git repository and works with almost any model through your own API key, so you pay only model costs [25]. On its own Aider polyglot benchmark, GPT-5 at high reasoning leads at 88.0 percent, though that leaderboard was last updated on 20 November 2025 and predates every 2026 model in this guide [24]. Aider is the go-to for developers who want full control and git-native, diff-based edits.

## Best AI code editors (IDE flow)

If you want AI inside a familiar editor with inline completions, chat, and an in-editor agent, these lead.

### 6. Cursor (Anysphere, now part of SpaceX)

Best for: the best all-around in-editor experience.

[Cursor](https://aiwiki.ai/wiki/cursor) is an AI-first code editor from [Anysphere](https://aiwiki.ai/wiki/anysphere), built on the VS Code codebase [48]. On 14 August 2026 Cursor announced that SpaceX had completed its acquisition of the company, following a model-training partnership with SpaceXAI announced in April [45]. Cursor now co-develops Grok models with SpaceXAI: it released Grok 4.6 "together with SpaceXAI" on 12 August and calls Grok 4.7 "our most capable model for long-running coding and knowledge work" [47][64]. Its own Composer 2.5 model (18 May 2026) is built on [Moonshot](https://aiwiki.ai/wiki/moonshot_ai)'s open [Kimi K2.5](https://aiwiki.ai/wiki/kimi_k2_5) checkpoint and costs $0.50 input and $2.50 output per million tokens, or $3 and $15 for the default fast variant [15]. vals.ai measured Composer 2.5 at 79.6 percent on SWE-bench Verified [3].

The Pro, Pro+, and Ultra plans have two usage pools: a "Cursor Models" pool with more included usage for Grok 4.7, Grok 4.6, Grok 4.5, and Composer 2.5, and an "Other Models" pool for third-party models such as Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5, Gemini 3.1 Pro, Gemini 3.8 Flash, and the GPT-5.6 models, charged at API rates [44]. Pricing: Hobby free, Pro $20 per month, Pro+ $60, Ultra $200, and Teams $40 per user (Standard) or $120 (Premium) [14][44]. Cursor's command-line agent, Cursor CLI, scored 79.3 percent with Grok 4.5 on the final Terminal-Bench 2.1 board but has no Terminal-Bench 4.0 entry [1][31].

### 7. GitHub Copilot (GitHub)

Best for: the widest rollout across editors and the safest enterprise choice.

[GitHub Copilot](https://aiwiki.ai/wiki/github_copilot), which GitHub calls "the world's most widely adopted AI developer tool," runs in VS Code, Visual Studio, JetBrains IDEs, Xcode, Neovim, Eclipse, Zed, and other editors, with chat, agent mode, a cloud agent, code review, and a CLI [16]. It is multi-model: its supported list includes GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, the GPT-5.6 models, Claude Fable 5.1, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5, Gemini 3.8 Flash, and Grok 4.7, and an Auto mode picks from a narrower subset (including GPT-6 Astra, the GPT-5.6 models, Claude Opus 5, and Claude Sonnet 5, but not Grok 4.7, Fable 5.1, Opus 5.5, or GPT-6 Sol), with GPT-5.3-Codex as the long-term-support fallback [18][49]. Paid plans also include access to third-party agents, including Claude Code and Codex [16]. On 1 June 2026 GitHub replaced premium requests with GitHub AI Credits consumed at each model's API rates [17]. Plans: Free; Pro $10 per month ($15 of credits); Pro+ $39 ($70 of credits); Max $100 ($200 of credits); Business $19 per user; Enterprise $39 per user [16]. GitHub does not publish its own SWE-bench figure, so benchmark quality tracks the model you pick [3].

### 8. Devin Desktop, formerly Windsurf (Cognition)

Best for: managing many local and cloud agents from a full IDE.

[Windsurf](https://aiwiki.ai/wiki/windsurf) was an AI IDE built around the Cascade agent. [Cognition](https://aiwiki.ai/wiki/cognition_ai), the maker of Devin, signed a definitive agreement to acquire Windsurf in July 2025, and on 2 June 2026 introduced Devin Desktop as "the next generation of Windsurf," with an agent command center as the default surface and full backward compatibility with Windsurf; windsurf.com now redirects to Devin Desktop [19][50]. On 10 September 2026 Cognition released SWE-2, post-trained from [Kimi K3](https://aiwiki.ai/wiki/kimi_k3), and calls it its most advanced coding model; Cognition reports 50.0 percent on FrontierCode 1.1 Main and 27.3 percent on Terminal-Bench 4, its own measurements [51]. SWE-2 follows SWE-1.7 and SWE-1.6; Cognition served SWE-1.6 at up to 950 tokens per second [51][21]. The free "SWE-2 Free" option is included in Devin Desktop and CLI through 10 October 2026 on the Pro plan, and frontier models from OpenAI, Anthropic, Google, and SpaceXAI are also selectable [20]. Pricing: Free, Pro $20 per month, Max $200 per month, Teams $80 per month plus $40 per full developer seat [20].

### 9. Cline (open source)

Best for: a transparent, open-source agent inside your existing IDE.

[Cline](https://aiwiki.ai/wiki/cline) is a free, open-source ([Apache-2.0](https://aiwiki.ai/wiki/apache_license)) autonomous coding agent available as an SDK, a VS Code extension, a CLI, and a JetBrains plugin (the JetBrains plugin itself is not open source) [26]. It is not tied to one provider: it works with Anthropic, OpenRouter, and many other model sources through your own key [26]. Because it is model-agnostic, its benchmark ceiling is whatever model you attach. It is the pick for developers who want an auditable agent without leaving their editor or paying a subscription.

## Best for enterprise and autonomous cloud engineering

### 10. Devin (Cognition)

Best for: fully autonomous, delegate-and-forget engineering at team scale.

[Devin](https://aiwiki.ai/wiki/devin) is [Cognition](https://aiwiki.ai/wiki/cognition_ai)'s autonomous software engineer: you assign it a task and it works in its own cloud environment, sold as "Devin Cloud" on the Pro plan and above [20]. SWE-2 is rolling out on Devin Web alongside Devin Desktop and the CLI [51]. Cognition has restructured pricing away from the old $500 per month plan: Free, Pro $20 per month (which adds Devin Cloud), Max $200 per month, Teams $80 per month plus $40 per seat, and custom Enterprise with VPC deployment [20]. Cognition does not publish a current SWE-bench Verified score for the Devin agent, so its cell is n/r.

### 11. Augment Code

Best for: large enterprise codebases needing deep context and compliance.

[Augment Code](https://aiwiki.ai/wiki/augment_code) is a coding agent with a CLI ("Auggie"), cloud agents, SOC 2 Type II compliance, and a "Context Engine" for large repositories [27][61]. Its documented results are 65.4 percent on SWE-bench Verified, set in March 2025 by combining Claude 3.7 Sonnet and OpenAI o1, and 51.8 percent on SWE-bench Pro for the Auggie CLI running [Claude Opus 4.5](https://aiwiki.ai/wiki/claude_opus_4_5), a result Augment published in February 2026 [28][61]. Pricing is flat rather than per seat: Standard $20 per month and Business $100 per month, each covering up to 50 seats with that amount of usage included, plus custom Enterprise [27]. Amp is another usage-based agent worth evaluating alongside it; it now runs its agents on remote machines it calls "orbs" [53].

## Open-weight coding models

Several open-weight models now score near the closed frontier on the saturated SWE-bench Verified, though they trail it on Terminal-Bench 4.0. All of them can be used from bring-your-own-model tools such as Aider and Cline, and Claude Code has a leaderboard entry with GLM-5.3 at 41.8 percent [31].

| Model | Weights license | SWE-bench Verified (vals.ai) | Terminal-Bench 4.0 (Artificial Analysis) |
|-------|-----------------|------------------------------|------------------------------------------|
| [DeepSeek V4](https://aiwiki.ai/wiki/deepseek_v4) Pro 0813 | [MIT](https://aiwiki.ai/wiki/mit_license) [57] | 96.4 [3] | n/t |
| [GLM-5.3](https://aiwiki.ai/wiki/glm_5_3) ([Zhipu AI](https://aiwiki.ai/wiki/zhipu_ai)) | custom GLM-5.3 license [58] | 95.4 [3] | 41.9 [41] |
| [Kimi K3](https://aiwiki.ai/wiki/kimi_k3) ([Moonshot AI](https://aiwiki.ai/wiki/moonshot_ai)) | custom Kimi K3 license [59] | 93.4 [3] | 12.6 [41] |
| [Qwen3.8](https://aiwiki.ai/wiki/qwen3_8) 27B | [Apache-2.0](https://aiwiki.ai/wiki/apache_license) [60] | 86.0 [3] | 5.6 [41] |

n/t: not tested or not listed by that evaluator. Cognition's SWE-2 is post-trained from Kimi K3, and Cursor's Composer 2.5 from Kimi K2.5 [51][15].

## Underlying models and API prices

The table below lists the models these tools run and their published API prices, so you can estimate real usage cost. Prices are USD per 1,000,000 tokens (standard tier, short context), verified 23 September 2026.

| Model | Developer | Input $/1M | Output $/1M | SWE-bench Verified (vals.ai) | Terminal-Bench 4.0 (Artificial Analysis) |
|-------|-----------|------------|-------------|------------------------------|------------------------------------------|
| [Claude Opus 5.5](https://aiwiki.ai/wiki/claude_opus_5_5) | [Anthropic](https://aiwiki.ai/wiki/anthropic) | 4 | 20 | n/t | 59.6 [41] |
| [Claude Fable 5.1](https://aiwiki.ai/wiki/claude_fable_5_1) | Anthropic | 10 | 50 | n/t | 55.1 [41] |
| [Claude Opus 5](https://aiwiki.ai/wiki/claude_opus_5) | Anthropic | 5 | 25 | 97.0 [3] | 49.0 [41] |
| [Claude Fable 5](https://aiwiki.ai/wiki/claude_fable_5) | Anthropic | 10 | 50 | 95.0 [3] | n/t |
| [Claude Opus 4.8](https://aiwiki.ai/wiki/claude_opus_4_8) | Anthropic | 5 | 25 | 88.6 [3] | n/t |
| [Claude Sonnet 5](https://aiwiki.ai/wiki/claude_sonnet_5) | Anthropic | 2 | 10 | 79.6 [3] | n/t |
| [GPT-6 Astra](https://aiwiki.ai/wiki/gpt_6_astra) | [OpenAI](https://aiwiki.ai/wiki/openai) | 10 | 50 | n/t | 59.6 [41] |
| [GPT-6 Sol](https://aiwiki.ai/wiki/gpt_6_sol) | OpenAI | 2 | 10 | n/t | 43.9 [41] |
| [GPT-5.6](https://aiwiki.ai/wiki/gpt_5_6) Sol | OpenAI | 4 (promotional) | 20 (promotional) | 96.2 [3] | 39.9 [41] |
| [GPT-5.5](https://aiwiki.ai/wiki/gpt-5.5) | OpenAI | 5 | 30 | 82.6 [3] | n/t |
| [Grok 4.7](https://aiwiki.ai/wiki/grok_4_7) | [SpaceXAI](https://aiwiki.ai/wiki/xai) | 2 | 6 | n/t | 25.8 [41] |
| [Gemini 3.1 Pro](https://aiwiki.ai/wiki/gemini_3_1_pro) (preview) | Google | 2 | 12 | 78.8 [3] | n/t |
| [Gemini 3.8 Flash](https://aiwiki.ai/wiki/gemini_3_8_flash) | Google | 0.75 (through 31 Dec 2026) | 3.75 (through 31 Dec 2026) | 80.0 [3] | 19.7 [41] |
| [Composer 2.5](https://aiwiki.ai/wiki/composer_2_5) | [Anysphere](https://aiwiki.ai/wiki/anysphere) | 0.50 | 2.50 | 79.6 [3] | n/t |

Anthropic prices are from its pricing page [5]; Claude Sonnet 5 launched at introductory pricing of $2 input and $10 output through 31 August 2026, and Anthropic then made that the standard price and cancelled the planned increase to $3 and $15 [5]. OpenAI prices are from its API pricing page and changelog; GPT-5.6 Sol's promotional pricing runs at least through 21 November 2026 [12][40]. Grok 4.7 prices are from SpaceXAI [46]; Gemini prices apply to prompts up to 200K tokens [23]; Composer 2.5 also has a fast variant at $3 input and $15 output [15]. Anthropic's own figure for Sonnet 5 on SWE-bench Verified, compiled by llm-stats, is 85.2 percent, higher than vals.ai's 79.6 [2][3]. Gemini 3.8 Flash's price rises to $1.50 input and $7.50 output from 1 January 2027 [23].

## SWE-bench Verified vs Terminal-Bench: how to read the scores

Two benchmarks matter here, and they measure different things. [SWE-bench Verified](https://aiwiki.ai/wiki/swe_bench_verified) is a set of 500 tasks drawn from real GitHub issues, each checked by running unit tests against the model's patch [3]. It is now close to saturated: on vals.ai's single-harness runs, seven models reach 95 percent or better and the leader, Claude Opus 5, scores 97.0 percent, so vals.ai no longer runs it on new releases; its results were last updated on 1 September 2026 [3]. Aggregators such as llm-stats mostly list vendors' self-reported Verified scores (116 self-reported results and no independently verified ones as of 23 September 2026), which can differ from independent runs [2][3].

[Terminal-Bench](https://aiwiki.ai/wiki/terminal_bench) scores an agent-plus-model combination on command-line tasks, so it captures how good the harness is, not only the model [31]. Watch three traps. First, versions are not comparable. Terminal-Bench 2.1 (6 May 2026) fixed 28 of the 89 tasks in 2.0, and most agent-model pairs scored higher on it than on 2.0, not lower [33]. Terminal-Bench 3.0 (30 July 2026) replaced it with 74 harder tasks [34], and Terminal-Bench 4.0 (28 August 2026) recalibrated task resources, fixed 19 tasks, and removed eight, including saturated ones [32]. Top scores fell from the 80s on 2.1 to under 60 percent on 4.0 [1][31]. Second, the public board itself changes: on 9 July 2026 its Terminal-Bench 2.1 table put Codex with GPT-5.5 first at 83.4 percent ahead of Claude Code with Fable 5 at 83.1 percent, but by 18 July a revised table had Claude Code with Fable 5 first at 83.8 percent and Codex with GPT-5.5 at 83.1 percent [56][63]. Third, vendor-reported numbers use their own setups. Anthropic's Opus 4.8 launch reported 74.6 percent on Terminal-Bench 2.1 with the public Terminus-2 harness, against 78.9 percent for Claude Code with Opus 4.8 on the public board [6][1]; Anthropic's 66.4 percent for Opus 5.5 on 4.0 is likewise its own run [35]. [Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis) runs every model through one harness (mini-SWE-agent for Terminal-Bench 4.0, averaging three runs per task), which is the most apples-to-apples model comparison: it puts Opus 5.5 and GPT-6 Astra at 59.6 percent, Fable 5.1 at 55.1 percent, and GPT-6 Sol at 43.9 percent [41]. On its older Terminal-Bench 2.1 runs, captured on 16 July 2026, it placed GPT-5.5 at 84.3 percent and Opus 4.8 at 84.6 percent [29]. For a tools comparison, the official leaderboard's agent-plus-model entries remain the most direct measure.

## Which AI coding assistant should you choose?

- Choose Claude Code if you want the highest ceiling on hard, autonomous, multi-file tasks and want Opus 5.5 by default with Fable 5.1 available.
- Choose OpenAI Codex if you already pay for ChatGPT, want the top Terminal-Bench 4.0 score at a lower run cost, or want an open-source CLI.
- Choose Cursor if you want the best day-to-day in-editor flow, with Grok and Composer models on a separate, larger usage pool.
- Choose GitHub Copilot if you need one tool across many editors, or you are buying for a team or enterprise.
- Choose Aider or Cline if you want free, open-source tooling and prefer to bring your own model, including open-weight models such as GLM-5.3 or Kimi K3.
- Choose Devin, Devin Desktop, or Augment Code if you want to hand off whole tasks or need enterprise-grade context and compliance.
- If you relied on Gemini CLI's free tier, move to Antigravity CLI or a paid Gemini API key.

## References

1. Terminal-Bench 2.1 leaderboard, tbench.ai. Wayback Machine capture of 26 August 2026 (the live URL now redirects to the 4.0 board). https://web.archive.org/web/20260826020311/https://www.tbench.ai/leaderboard/terminal-bench/2.1
2. SWE-bench Verified leaderboard, llm-stats.com (self-reported results; last updated 23 September 2026). https://llm-stats.com/benchmarks/swe-bench-verified
3. SWE-bench Verified benchmark (archived benchmark, updated 1 September 2026), vals.ai. https://www.vals.ai/benchmarks/swebench
4. Anthropic, Claude Code model configuration and defaults. https://code.claude.com/docs/en/model-config
5. Anthropic, Claude model and plan pricing. https://platform.claude.com/docs/en/about-claude/pricing and https://claude.com/pricing
6. Anthropic, Introducing Claude Opus 4.8. https://www.anthropic.com/news/claude-opus-4-8
7. Anthropic, Redeploying Claude Fable 5. https://www.anthropic.com/news/redeploying-fable-5
8. Google DeepMind, Gemini 3.1 Pro model card. https://deepmind.google/models/model-cards/gemini-3-1-pro/
9. OpenAI, Codex models. https://developers.openai.com/codex/models
10. OpenAI, Codex pricing. https://developers.openai.com/codex/pricing
11. OpenAI, Introducing GPT-5.5. https://openai.com/index/introducing-gpt-5-5/
12. OpenAI, API pricing. https://developers.openai.com/api/docs/pricing
13. OpenAI Codex CLI, GitHub (Apache-2.0). https://github.com/openai/codex
14. Cursor pricing. https://cursor.com/pricing
15. Cursor, Composer 2.5 (18 May 2026). https://cursor.com/blog/composer-2-5
16. GitHub Copilot plans. https://github.com/features/copilot/plans
17. GitHub, GitHub Copilot is moving to usage-based billing (27 April 2026). https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/
18. GitHub Copilot supported AI models. https://docs.github.com/en/copilot/reference/ai-models/supported-models
19. Cognition, Cognition's acquisition of Windsurf (14 July 2025). https://cognition.com/blog/windsurf
20. Devin, Plans and Pricing. https://devin.ai/pricing
21. Cognition, An Early Preview of SWE-1.6 and Research Update. https://cognition.com/blog/swe-1-6-preview
22. Gemini CLI, GitHub (Apache-2.0). https://github.com/google-gemini/gemini-cli
23. Google, Gemini API pricing. https://ai.google.dev/gemini-api/docs/pricing
24. Aider LLM leaderboards. https://aider.chat/docs/leaderboards/
25. Aider, GitHub (Apache-2.0). https://github.com/Aider-AI/aider
26. Cline, GitHub (Apache-2.0). https://github.com/cline/cline
27. Augment Code pricing. https://www.augmentcode.com/pricing
28. Augment Code, #1 open-source agent on SWE-Bench Verified by combining Claude 3.7 and O1 (31 March 2025). https://www.augmentcode.com/blog/1-open-source-agent-on-swe-bench-verified-by-combining-claude-3-7-and-o1
29. Artificial Analysis, Terminal-Bench v2.1 (Wayback Machine capture of 16 July 2026). https://web.archive.org/web/20260716053902/https://artificialanalysis.ai/evaluations/terminalbench-v2-1
30. Google Cloud, Gemini Code Assist Standard and Enterprise overview (Antigravity migration note). https://docs.cloud.google.com/gemini/docs/codeassist/overview
31. Terminal-Bench 4.0 leaderboard, tbench.ai (accessed 23 September 2026). https://www.tbench.ai/leaderboard/terminal-bench/4.0
32. Ryan Marten, Terminal-Bench 4.0, Terminal-Bench blog (28 August 2026). https://www.tbench.ai/blog/terminal-bench-4-0
33. Terminal-Bench 2.1, Terminal-Bench blog (6 May 2026). https://www.tbench.ai/blog/terminal-bench-2-1
34. Ryan Marten, Alex Shaw, and Andy Konwinski, Terminal-Bench 3.0 measures agent abilities at the frontier, Terminal-Bench blog (30 July 2026). https://www.tbench.ai/blog/terminal-bench-3-0
35. Anthropic, Introducing Claude Opus 5.5 (22 September 2026). https://www.anthropic.com/claude-opus-5-5
36. Anthropic, Introducing Claude Fable 5.1 and Claude Mythos 5.1 (September 2026). https://www.anthropic.com/claude-fable-and-mythos-5-1
37. Anthropic, Introducing Claude Opus 5 (24 July 2026). https://www.anthropic.com/news/claude-opus-5
38. OpenAI, GPT-6 Astra: A new generation of intelligence (3 September 2026). https://openai.com/index/gpt-6-astra/
39. OpenAI, Introducing GPT-6 Sol and Luna (22 September 2026). https://openai.com/index/introducing-gpt-6-sol-and-luna/
40. OpenAI, API changelog. https://developers.openai.com/api/docs/changelog
41. Artificial Analysis, Terminal-Bench 4.0 Benchmark Leaderboard (accessed 23 September 2026). https://artificialanalysis.ai/evaluations/terminalbench-v4-0
42. Dmitry Lyalin and Taylor Mullen, An important update: Transitioning Gemini CLI to Antigravity CLI, Google Developers Blog (19 May 2026). https://developers.googleblog.com/an-important-update-transitioning-gemini-cli-to-antigravity-cli/
43. Google, Gemini Code Assist consumer accounts (deprecation notice). https://developers.google.com/gemini-code-assist/docs/deprecations/code-assist-individuals
44. Cursor Docs, Models and Pricing. https://cursor.com/docs/models
45. Cursor, Cursor is now a part of SpaceX (14 August 2026). https://cursor.com/blog/joining-spacex
46. SpaceXAI, Introducing Grok 4.7 (21 September 2026). https://x.ai/news/grok-4-7 (launch-day capture with the original 38.0% figure: https://web.archive.org/web/20260921182849/https://x.ai/news/grok-4-7)
47. Cursor, Introducing Grok 4.6 (August 2026). https://cursor.com/blog/grok-4-6
48. Cursor Docs, VS Code Migration. https://cursor.com/docs/configuration/migrations/vscode
49. GitHub Changelog, Grok 4.7 is now available in GitHub Copilot (21 September 2026). https://github.blog/changelog/2026-09-21-grok-4-7-is-now-available-in-github-copilot/
50. Scott Wu and Jeff Wang, Introducing Devin Desktop, Cognition (2 June 2026). https://cognition.com/blog/introducing-devin-desktop
51. Cognition, Introducing SWE-2: Pushing the Pareto Frontier (10 September 2026). https://cognition.com/blog/swe-2
52. Anthropic, Claude Code GitHub repository. https://github.com/anthropics/claude-code
53. Amp, Coding agent and dev environment built for the frontier. https://ampcode.com
54. Anthropic, Claude Code overview. https://code.claude.com/docs/en/overview
55. Anthropic, Claude Code license notice ("All rights reserved"). https://github.com/anthropics/claude-code/blob/main/LICENSE.md
56. Terminal-Bench 2.1 leaderboard, tbench.ai. Wayback Machine capture of 9 July 2026. https://web.archive.org/web/20260709181004/https://www.tbench.ai/leaderboard/terminal-bench/2.1
57. DeepSeek, DeepSeek-V4-Pro-0813, Hugging Face (MIT license). https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
58. Zhipu AI (Z.ai), GLM-5.3, Hugging Face (GLM-5.3 license). https://huggingface.co/zai-org/GLM-5.3
59. Moonshot AI, Kimi-K3, Hugging Face (Kimi K3 license). https://huggingface.co/moonshotai/Kimi-K3
60. Qwen, Qwen3.8-27B, Hugging Face (Apache-2.0). https://huggingface.co/Qwen/Qwen3.8-27B
61. Augment Code, Auggie tops SWE-Bench Pro (4 February 2026). https://www.augmentcode.com/blog/auggie-tops-swe-bench-pro
62. Anthropic, What is the Max plan? Claude Help Center. https://support.claude.com/en/articles/11049741-what-is-the-max-plan
63. Terminal-Bench 2.1 leaderboard, tbench.ai. Wayback Machine capture of 18 July 2026. https://web.archive.org/web/20260718002441/https://www.tbench.ai/leaderboard/terminal-bench/2.1
64. Cursor, Introducing Grok 4.7 (21 September 2026). https://cursor.com/blog/grok-4-7

