Citation and evidence

Best AI Coding Assistants

24 min full readUpdated 64 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI AgentsAI Code GenerationDeveloper ToolsLarge Language Models

Cite this article

As of 23 September 2026, the two tools at the top of the official Terminal-Bench 4.0 leaderboard are OpenAI's Codex and Anthropic's Claude Code. Codex running GPT-6 Astra at max effort scores 58.2 percent, and Claude Code running Claude Fable 5.1 at max effort scores 57.9 percent, a gap well inside the board's roughly 3-point confidence intervals [31]. Neither Anthropic's Claude Opus 5.5, released on 22 September 2026 and now the default model in Claude Code, nor OpenAI's GPT-6 Sol, released the same day, had a public leaderboard entry yet; Anthropic reports 66.4 percent for Opus 5.5 on its own reproduction of the benchmark, and the independent evaluator Artificial Analysis measures Opus 5.5 and GPT-6 Astra level at 59.6 percent with a single shared harness [35][41]. For in-editor (IDE) coding, Cursor, now owned by SpaceX, and GitHub Copilot remain the main choices, and Cognition has folded Windsurf into its Devin Desktop editor alongside the Devin cloud agent [45][50].

Quick verdict: which AI coding assistant is best?

  • Highest ceiling on hard, autonomous, multi-file work: Claude Code (Opus 5.5 by default, Claude Fable 5.1 on request) or OpenAI Codex (GPT-6 Astra). They hold the top five places on Terminal-Bench 4.0 between them [31].
  • Best value at the top of the board: OpenAI Codex with GPT-6 Astra, whose max-effort leaderboard run cost about $3.3k against about $6.2k for Claude Code with Fable 5.1 at max effort, at essentially the same score [31]. Codex is included in every ChatGPT plan, from Free upward [10].
  • Best AI IDE (in-editor flow): Cursor, which pairs its first-party Grok and Composer models with frontier models from Anthropic, OpenAI, and Google [44].
  • Best for the widest team and enterprise rollout: GitHub Copilot, which runs in VS Code, Visual Studio, JetBrains IDEs, Xcode, Neovim, Eclipse, Zed, and other editors and offers models from OpenAI, Anthropic, Google, and xAI [16][18].
  • Best fully autonomous cloud engineer: Devin, with Cognition's new SWE-2 model [51].
  • Best open-source and model-agnostic: Aider (terminal) and Cline (IDE and CLI) [25][26].
  • Free terminal option: Gemini CLI no longer is one. Google ended free personal-account access on 18 June 2026 and points individual users to Antigravity CLI [42][43].

Summary comparison table

ToolDeveloperDefault / notable modelsSWE-bench Verified (vals.ai)Terminal-Bench 4.0 (tbench.ai)Price (USD)Type / access
Claude CodeAnthropicOpus 5.5 default; Fable 5.1 selectable97.0 (Opus 5) / 95.0 (Fable 5) [3]57.9 (Fable 5.1) / 53.9 (Opus 5) [31]Pro $20, Max $100 or $200/mo, or API [5][62]Terminal agent + IDE, desktop, web; proprietary [54][55]
OpenAI Codex CLIOpenAIGPT-6 Sol recommended; GPT-6 Astra96.2 (GPT-5.6 Sol) [3]58.2 (GPT-6 Astra) / 37.3 (GPT-5.6 Sol) [31]Free; Go $8, Plus $20, Pro from $100, Business $20/user; or API [10]Terminal agent + IDE + desktop app + cloud; CLI open source (Apache-2.0) [13]
CursorAnysphere (SpaceX)Grok 4.7, Composer 2.5 + frontier79.6 (Composer 2.5) [3]n/rHobby free; Pro $20, Pro+ $60, Ultra $200/mo; Teams $40/user [14]AI IDE (VS Code codebase), CLI, cloud agents [48]
GitHub CopilotGitHubMulti-model (Auto; GPT-6, Claude, Gemini, Grok)up to 97.0 (via Opus 5) [3]n/rFree; Pro $10, Pro+ $39, Max $100/mo; Business $19, Enterprise $39/user [16]IDE ext + agent + CLI
Devin Desktop (formerly Windsurf)CognitionSWE-2 + frontiern/rn/rFree; Pro $20, Max $200/mo; Teams $80 + $40/seat [20]AI IDE with agent manager
DevinCognitionSWE-2 + frontiern/rn/rCloud agents from Pro $20/mo [20]Autonomous cloud agent
Grok BuildSpaceXAIGrok 4.7n/r37.6 (Grok 4.7) [31]Free to try; Grok API [46]Terminal agent
Gemini CLIGoogleGemini 3 models78.8 (Gemini 3.1 Pro) [3]n/rPaid Gemini API key or Code Assist Standard/Enterprise [42]Terminal agent, open source (Apache-2.0) [22]
AiderAider-AI (open source)Bring your own modelinherits modeln/rFree, pay model APITerminal pair programmer, open source (Apache-2.0) [25]
ClineCline Bot Inc. (open source)Bring your own modelinherits modeln/rFree, pay model APISDK, IDE ext, CLI, open source (Apache-2.0) [26]
Augment CodeAugmentContext Engine + LLMs (Opus 4.5 in its SWE-bench Pro run)n/r (self-reports 65.4, March 2025 [28])n/rStandard $20/mo, Business $100/mo (flat, up to 50 seats), Enterprise custom [27]CLI + cloud agents

Expanded article table

Last verified: 23 September 2026. Prices in USD. The SWE-bench Verified column uses vals.ai's independent results with a single bash-only harness; vals.ai has archived that benchmark as saturated (its results were last updated on 1 September 2026) and does not test newer models such as Opus 5.5, Fable 5.1, GPT-6 Astra, or GPT-6 Sol [3]. The Terminal-Bench 4.0 column is a tool-plus-model score from the official leaderboard, best result across reasoning-effort settings [31]. An n/r cell means no current primary source reports that figure. Bring-your-own-model tools inherit the score of whichever model you attach.

Best terminal coding agents (CLI)

Command-line agents run in your terminal, read and edit the whole repository, run tests, and iterate on failures. They hold every place on the Terminal-Bench 4.0 leaderboard: the board lists entries for Codex, Claude Code, Grok Build, and the minimal mini-SWE-agent harness [31].

1. Claude Code (Anthropic)

Best for: complex, autonomous, multi-file engineering where correctness matters most.

Claude Code is Anthropic's agentic coding tool. It runs in the terminal, in VS Code and JetBrains IDEs, in a desktop app, and on the web [54]. It is proprietary software: its GitHub repository [52] carries an "All rights reserved" notice [55]. Since version 2.1.280 the default model on Pro, Max, Team, Enterprise, and the Anthropic API is Claude Opus 5.5; before that, Pro and Team Standard defaulted to Claude Sonnet 5 and the other plans to Claude Opus 5 [4]. Claude Fable 5.1 and Fable 5 are selectable but are not the default on any plan [4]. Fable 5 was briefly withdrawn after US export controls were applied on 12 June 2026, returned on 1 July, and moved to usage credits on subscription plans after 7 July [7]. On claude.com's plan table, Fable runs on usage credits on Pro and within 50 percent of weekly limits on Max [5].

On the public Terminal-Bench 4.0 board, Claude Code scores 57.9 percent with Fable 5.1 and 53.9 percent with Opus 5 [31]. Opus 5.5 was not yet listed on 23 September; Anthropic reports 66.4 percent for it at xhigh effort on a setup that reproduces the leaderboard's Opus 5 result within noise, and Artificial Analysis measures it at 59.6 percent [35][41]. Anthropic says Opus 5.5 performs at the level of Fable 5.1 on most work and costs 40 percent less to run than Opus 5 [35]. Plans: Pro $20 per month ($17 with annual billing), Max $100 per month (5x Pro usage) or $200 per month (20x), plus API pay-as-you-go [5][62]. API token prices per million are $4 input and $20 output for Opus 5.5, $10 and $50 for Fable 5.1, and $5 and $25 for Opus 5, which launched on 24 July 2026 [5][37]. Fable 5.1 kept Fable 5's $10 and $50 pricing [36]. With Opus 5.5, Anthropic also raised five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans [35].

2. OpenAI Codex CLI (OpenAI)

Best for: top-of-board agentic coding at lower cost, and anyone already paying for ChatGPT.

The OpenAI Codex CLI is OpenAI's open-source terminal agent, released under the Apache-2.0 license, and part of a Codex product that also runs in the ChatGPT desktop app, IDE extensions, and the cloud [13][9]. OpenAI released GPT-6 Astra on 3 September 2026 and GPT-6 Sol and GPT-6 Luna on 22 September 2026 [38][39][40]. Codex's documentation recommends GPT-6 Sol for complex coding and agentic work and GPT-6 Astra for the hardest end-to-end tasks; Astra is available in the Codex CLI and IDE extension but not in Codex cloud [9]. GPT-5.5 retires from Codex for ChatGPT sign-ins on 14 October 2026 [9].

Codex with GPT-6 Astra holds first place on Terminal-Bench 4.0 at 58.2 percent (max effort), and ties Claude Code with Fable 5.1 at 57.9 percent at high and xhigh effort [31]. The previous-generation GPT-5.6 Sol, released on 9 July 2026, scores 37.3 percent in Codex [31][40]. OpenAI's own GPT-5.5 launch in April reported 82.7 percent on the older Terminal-Bench 2.0 and 58.6 percent on SWE-bench Pro, without a SWE-bench Verified figure; vals.ai later measured GPT-5.5 at 82.6 percent on Verified [11][3]. Codex is included in ChatGPT Free, Go ($8), Plus ($20), Pro (from $100 per month), and Business ($20 per user per month billed annually) [10]. API prices per million tokens are $10 input and $50 output for GPT-6 Astra, $2 and $10 for GPT-6 Sol, and $5 and $30 for GPT-5.5 [12][38][39]. OpenAI cut Sol and Luna prices by half compared with GPT-5.6's promotional pricing [39].

3. Grok Build (SpaceXAI)

Best for: developers already using Grok models, and a free way to try them in the terminal.

Grok Build is SpaceXAI's coding agent, installed from x.ai/build [46]. With Grok 4.7, released on 21 September 2026, it scores 37.6 percent on Terminal-Bench 4.0 (over 330 trials; the leaderboard marks only its cost figure as partial, covering 324 of 330 trials), up from 20.3 percent with Grok 4.6 [31][46]. SpaceXAI's launch page first gave Grok 4.7 38.0 percent on Terminal-Bench 4.0 and later changed the figure to 37.6 percent, which matches the leaderboard [46]. Grok 4.7 costs $2 per million input tokens and $6 per million output tokens on the Grok API, and SpaceXAI invites users to try it in Grok Build for free [46].

4. Gemini CLI and Antigravity CLI (Google)

Best for: teams that already hold Gemini Code Assist Standard or Enterprise licenses or paid Gemini API keys.

Gemini CLI is Google's open-source (Apache-2.0) terminal agent [22]. Its main draw used to be a free tier through a personal Google account (60 requests per minute and 1,000 per day) [22], but on 19 May 2026 Google announced that it was unifying its terminal tools into Antigravity and its new Antigravity CLI [42]. From 18 June 2026, Gemini CLI stopped serving requests for Gemini Code Assist for individuals, Google AI Pro, and Google AI Ultra, and the "Login with Google" option no longer works; affected users are directed to Antigravity and Antigravity CLI [43][30]. Gemini CLI remains available through paid Gemini API keys and through Code Assist Standard and Enterprise licenses [42]. Google's model card gives Gemini 3.1 Pro 80.6 percent on SWE-bench Verified, while vals.ai's harness measures the preview at 78.8 percent [8][3]. On the last Terminal-Bench 2.1 board, Gemini CLI with Gemini 3.1 Pro scored 65.8 percent; it has no Terminal-Bench 4.0 entry, and Google's Gemini 3.8 Flash appears there only with the mini-SWE-agent harness, at 19.1 percent [1][31]. Gemini API prices are $2 input and $12 output per million tokens for Gemini 3.1 Pro Preview (prompts up to 200K tokens) and $0.75 and $3.75 for Gemini 3.8 Flash through 31 December 2026 [23].

5. Aider (open source)

Best for: a lightweight, scriptable, model-agnostic pair programmer in the terminal.

Aider is a free, open-source (Apache-2.0) terminal tool that edits your git repository and works with almost any model through your own API key, so you pay only model costs [25]. On its own Aider polyglot benchmark, GPT-5 at high reasoning leads at 88.0 percent, though that leaderboard was last updated on 20 November 2025 and predates every 2026 model in this guide [24]. Aider is the go-to for developers who want full control and git-native, diff-based edits.

Best AI code editors (IDE flow)

If you want AI inside a familiar editor with inline completions, chat, and an in-editor agent, these lead.

6. Cursor (Anysphere, now part of SpaceX)

Best for: the best all-around in-editor experience.

Cursor is an AI-first code editor from Anysphere, built on the VS Code codebase [48]. On 14 August 2026 Cursor announced that SpaceX had completed its acquisition of the company, following a model-training partnership with SpaceXAI announced in April [45]. Cursor now co-develops Grok models with SpaceXAI: it released Grok 4.6 "together with SpaceXAI" on 12 August and calls Grok 4.7 "our most capable model for long-running coding and knowledge work" [47][64]. Its own Composer 2.5 model (18 May 2026) is built on Moonshot's open Kimi K2.5 checkpoint and costs $0.50 input and $2.50 output per million tokens, or $3 and $15 for the default fast variant [15]. vals.ai measured Composer 2.5 at 79.6 percent on SWE-bench Verified [3].

The Pro, Pro+, and Ultra plans have two usage pools: a "Cursor Models" pool with more included usage for Grok 4.7, Grok 4.6, Grok 4.5, and Composer 2.5, and an "Other Models" pool for third-party models such as Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5, Gemini 3.1 Pro, Gemini 3.8 Flash, and the GPT-5.6 models, charged at API rates [44]. Pricing: Hobby free, Pro $20 per month, Pro+ $60, Ultra $200, and Teams $40 per user (Standard) or $120 (Premium) [14][44]. Cursor's command-line agent, Cursor CLI, scored 79.3 percent with Grok 4.5 on the final Terminal-Bench 2.1 board but has no Terminal-Bench 4.0 entry [1][31].

7. GitHub Copilot (GitHub)

Best for: the widest rollout across editors and the safest enterprise choice.

GitHub Copilot, which GitHub calls "the world's most widely adopted AI developer tool," runs in VS Code, Visual Studio, JetBrains IDEs, Xcode, Neovim, Eclipse, Zed, and other editors, with chat, agent mode, a cloud agent, code review, and a CLI [16]. It is multi-model: its supported list includes GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, the GPT-5.6 models, Claude Fable 5.1, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5, Gemini 3.8 Flash, and Grok 4.7, and an Auto mode picks from a narrower subset (including GPT-6 Astra, the GPT-5.6 models, Claude Opus 5, and Claude Sonnet 5, but not Grok 4.7, Fable 5.1, Opus 5.5, or GPT-6 Sol), with GPT-5.3-Codex as the long-term-support fallback [18][49]. Paid plans also include access to third-party agents, including Claude Code and Codex [16]. On 1 June 2026 GitHub replaced premium requests with GitHub AI Credits consumed at each model's API rates [17]. Plans: Free; Pro $10 per month ($15 of credits); Pro+ $39 ($70 of credits); Max $100 ($200 of credits); Business $19 per user; Enterprise $39 per user [16]. GitHub does not publish its own SWE-bench figure, so benchmark quality tracks the model you pick [3].

8. Devin Desktop, formerly Windsurf (Cognition)

Best for: managing many local and cloud agents from a full IDE.

Windsurf was an AI IDE built around the Cascade agent. Cognition, the maker of Devin, signed a definitive agreement to acquire Windsurf in July 2025, and on 2 June 2026 introduced Devin Desktop as "the next generation of Windsurf," with an agent command center as the default surface and full backward compatibility with Windsurf; windsurf.com now redirects to Devin Desktop [19][50]. On 10 September 2026 Cognition released SWE-2, post-trained from Kimi K3, and calls it its most advanced coding model; Cognition reports 50.0 percent on FrontierCode 1.1 Main and 27.3 percent on Terminal-Bench 4, its own measurements [51]. SWE-2 follows SWE-1.7 and SWE-1.6; Cognition served SWE-1.6 at up to 950 tokens per second [51][21]. The free "SWE-2 Free" option is included in Devin Desktop and CLI through 10 October 2026 on the Pro plan, and frontier models from OpenAI, Anthropic, Google, and SpaceXAI are also selectable [20]. Pricing: Free, Pro $20 per month, Max $200 per month, Teams $80 per month plus $40 per full developer seat [20].

9. Cline (open source)

Best for: a transparent, open-source agent inside your existing IDE.

Cline is a free, open-source (Apache-2.0) autonomous coding agent available as an SDK, a VS Code extension, a CLI, and a JetBrains plugin (the JetBrains plugin itself is not open source) [26]. It is not tied to one provider: it works with Anthropic, OpenRouter, and many other model sources through your own key [26]. Because it is model-agnostic, its benchmark ceiling is whatever model you attach. It is the pick for developers who want an auditable agent without leaving their editor or paying a subscription.

Best for enterprise and autonomous cloud engineering

10. Devin (Cognition)

Best for: fully autonomous, delegate-and-forget engineering at team scale.

Devin is Cognition's autonomous software engineer: you assign it a task and it works in its own cloud environment, sold as "Devin Cloud" on the Pro plan and above [20]. SWE-2 is rolling out on Devin Web alongside Devin Desktop and the CLI [51]. Cognition has restructured pricing away from the old $500 per month plan: Free, Pro $20 per month (which adds Devin Cloud), Max $200 per month, Teams $80 per month plus $40 per seat, and custom Enterprise with VPC deployment [20]. Cognition does not publish a current SWE-bench Verified score for the Devin agent, so its cell is n/r.

11. Augment Code

Best for: large enterprise codebases needing deep context and compliance.

Augment Code is a coding agent with a CLI ("Auggie"), cloud agents, SOC 2 Type II compliance, and a "Context Engine" for large repositories [27][61]. Its documented results are 65.4 percent on SWE-bench Verified, set in March 2025 by combining Claude 3.7 Sonnet and OpenAI o1, and 51.8 percent on SWE-bench Pro for the Auggie CLI running Claude Opus 4.5, a result Augment published in February 2026 [28][61]. Pricing is flat rather than per seat: Standard $20 per month and Business $100 per month, each covering up to 50 seats with that amount of usage included, plus custom Enterprise [27]. Amp is another usage-based agent worth evaluating alongside it; it now runs its agents on remote machines it calls "orbs" [53].

Open-weight coding models

Several open-weight models now score near the closed frontier on the saturated SWE-bench Verified, though they trail it on Terminal-Bench 4.0. All of them can be used from bring-your-own-model tools such as Aider and Cline, and Claude Code has a leaderboard entry with GLM-5.3 at 41.8 percent [31].

ModelWeights licenseSWE-bench Verified (vals.ai)Terminal-Bench 4.0 (Artificial Analysis)
DeepSeek V4 Pro 0813MIT [57]96.4 [3]n/t
GLM-5.3 (Zhipu AI)custom GLM-5.3 license [58]95.4 [3]41.9 [41]
Kimi K3 (Moonshot AI)custom Kimi K3 license [59]93.4 [3]12.6 [41]
Qwen3.8 27BApache-2.0 [60]86.0 [3]5.6 [41]

Expanded article table

n/t: not tested or not listed by that evaluator. Cognition's SWE-2 is post-trained from Kimi K3, and Cursor's Composer 2.5 from Kimi K2.5 [51][15].

Underlying models and API prices

The table below lists the models these tools run and their published API prices, so you can estimate real usage cost. Prices are USD per 1,000,000 tokens (standard tier, short context), verified 23 September 2026.

ModelDeveloperInput $/1MOutput $/1MSWE-bench Verified (vals.ai)Terminal-Bench 4.0 (Artificial Analysis)
Claude Opus 5.5Anthropic420n/t59.6 [41]
Claude Fable 5.1Anthropic1050n/t55.1 [41]
Claude Opus 5Anthropic52597.0 [3]49.0 [41]
Claude Fable 5Anthropic105095.0 [3]n/t
Claude Opus 4.8Anthropic52588.6 [3]n/t
Claude Sonnet 5Anthropic21079.6 [3]n/t
GPT-6 AstraOpenAI1050n/t59.6 [41]
GPT-6 SolOpenAI210n/t43.9 [41]
GPT-5.6 SolOpenAI4 (promotional)20 (promotional)96.2 [3]39.9 [41]
GPT-5.5OpenAI53082.6 [3]n/t
Grok 4.7SpaceXAI26n/t25.8 [41]
Gemini 3.1 Pro (preview)Google21278.8 [3]n/t
Gemini 3.8 FlashGoogle0.75 (through 31 Dec 2026)3.75 (through 31 Dec 2026)80.0 [3]19.7 [41]
Composer 2.5Anysphere0.502.5079.6 [3]n/t

Expanded article table

Anthropic prices are from its pricing page [5]; Claude Sonnet 5 launched at introductory pricing of $2 input and $10 output through 31 August 2026, and Anthropic then made that the standard price and cancelled the planned increase to $3 and $15 [5]. OpenAI prices are from its API pricing page and changelog; GPT-5.6 Sol's promotional pricing runs at least through 21 November 2026 [12][40]. Grok 4.7 prices are from SpaceXAI [46]; Gemini prices apply to prompts up to 200K tokens [23]; Composer 2.5 also has a fast variant at $3 input and $15 output [15]. Anthropic's own figure for Sonnet 5 on SWE-bench Verified, compiled by llm-stats, is 85.2 percent, higher than vals.ai's 79.6 [2][3]. Gemini 3.8 Flash's price rises to $1.50 input and $7.50 output from 1 January 2027 [23].

SWE-bench Verified vs Terminal-Bench: how to read the scores

Two benchmarks matter here, and they measure different things. SWE-bench Verified is a set of 500 tasks drawn from real GitHub issues, each checked by running unit tests against the model's patch [3]. It is now close to saturated: on vals.ai's single-harness runs, seven models reach 95 percent or better and the leader, Claude Opus 5, scores 97.0 percent, so vals.ai no longer runs it on new releases; its results were last updated on 1 September 2026 [3]. Aggregators such as llm-stats mostly list vendors' self-reported Verified scores (116 self-reported results and no independently verified ones as of 23 September 2026), which can differ from independent runs [2][3].

Terminal-Bench scores an agent-plus-model combination on command-line tasks, so it captures how good the harness is, not only the model [31]. Watch three traps. First, versions are not comparable. Terminal-Bench 2.1 (6 May 2026) fixed 28 of the 89 tasks in 2.0, and most agent-model pairs scored higher on it than on 2.0, not lower [33]. Terminal-Bench 3.0 (30 July 2026) replaced it with 74 harder tasks [34], and Terminal-Bench 4.0 (28 August 2026) recalibrated task resources, fixed 19 tasks, and removed eight, including saturated ones [32]. Top scores fell from the 80s on 2.1 to under 60 percent on 4.0 [1][31]. Second, the public board itself changes: on 9 July 2026 its Terminal-Bench 2.1 table put Codex with GPT-5.5 first at 83.4 percent ahead of Claude Code with Fable 5 at 83.1 percent, but by 18 July a revised table had Claude Code with Fable 5 first at 83.8 percent and Codex with GPT-5.5 at 83.1 percent [56][63]. Third, vendor-reported numbers use their own setups. Anthropic's Opus 4.8 launch reported 74.6 percent on Terminal-Bench 2.1 with the public Terminus-2 harness, against 78.9 percent for Claude Code with Opus 4.8 on the public board [6][1]; Anthropic's 66.4 percent for Opus 5.5 on 4.0 is likewise its own run [35]. Artificial Analysis runs every model through one harness (mini-SWE-agent for Terminal-Bench 4.0, averaging three runs per task), which is the most apples-to-apples model comparison: it puts Opus 5.5 and GPT-6 Astra at 59.6 percent, Fable 5.1 at 55.1 percent, and GPT-6 Sol at 43.9 percent [41]. On its older Terminal-Bench 2.1 runs, captured on 16 July 2026, it placed GPT-5.5 at 84.3 percent and Opus 4.8 at 84.6 percent [29]. For a tools comparison, the official leaderboard's agent-plus-model entries remain the most direct measure.

Which AI coding assistant should you choose?

  • Choose Claude Code if you want the highest ceiling on hard, autonomous, multi-file tasks and want Opus 5.5 by default with Fable 5.1 available.
  • Choose OpenAI Codex if you already pay for ChatGPT, want the top Terminal-Bench 4.0 score at a lower run cost, or want an open-source CLI.
  • Choose Cursor if you want the best day-to-day in-editor flow, with Grok and Composer models on a separate, larger usage pool.
  • Choose GitHub Copilot if you need one tool across many editors, or you are buying for a team or enterprise.
  • Choose Aider or Cline if you want free, open-source tooling and prefer to bring your own model, including open-weight models such as GLM-5.3 or Kimi K3.
  • Choose Devin, Devin Desktop, or Augment Code if you want to hand off whole tasks or need enterprise-grade context and compliance.
  • If you relied on Gemini CLI's free tier, move to Antigravity CLI or a paid Gemini API key.

References

  1. ^1 ^2 ^3 ^4Terminal-Bench 2.1 leaderboard, tbench.ai. Wayback Machine capture of 26 August 2026 (the live URL now redirects to the 4.0 board). web.archive.org/...2.1
  2. ^1 ^2SWE-bench Verified leaderboard, llm-stats.com (self-reported results; last updated 23 September 2026). llm-stats.com/...swe-bench-verified
  3. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27SWE-bench Verified benchmark (archived benchmark, updated 1 September 2026), vals.ai. vals.ai/...swebench
  4. ^1 ^2Anthropic, Claude Code model configuration and defaults. code.claude.com/...model-config
  5. ^1 ^2 ^3 ^4 ^5 ^6Anthropic, Claude model and plan pricing. platform.claude.com/...pricing and claude.com/pricing
  6. ^Anthropic, Introducing Claude Opus 4.8. anthropic.com/...claude-opus-4-8
  7. ^Anthropic, Redeploying Claude Fable 5. anthropic.com/...redeploying-fable-5
  8. ^Google DeepMind, Gemini 3.1 Pro model card. deepmind.google/...gemini-3-1-pro
  9. ^1 ^2 ^3OpenAI, Codex models. developers.openai.com/...models
  10. ^1 ^2 ^3OpenAI, Codex pricing. developers.openai.com/...pricing
  11. ^OpenAI, Introducing GPT-5.5. openai.com/...introducing-gpt-5-5
  12. ^1 ^2OpenAI, API pricing. developers.openai.com/...pricing
  13. ^1 ^2OpenAI Codex CLI, GitHub (Apache-2.0). github.com/...codex
  14. ^1 ^2Cursor pricing. cursor.com/pricing
  15. ^1 ^2 ^3Cursor, Composer 2.5 (18 May 2026). cursor.com/...composer-2-5
  16. ^1 ^2 ^3 ^4 ^5GitHub Copilot plans. github.com/...plans
  17. ^GitHub, GitHub Copilot is moving to usage-based billing (27 April 2026). github.blog/...ot-is-moving-to-usage-based-billing
  18. ^1 ^2GitHub Copilot supported AI models. docs.github.com/...supported-models
  19. ^Cognition, Cognition's acquisition of Windsurf (14 July 2025). cognition.com/...windsurf
  20. ^1 ^2 ^3 ^4 ^5 ^6Devin, Plans and Pricing. devin.ai/pricing
  21. ^Cognition, An Early Preview of SWE-1.6 and Research Update. cognition.com/...swe-1-6-preview
  22. ^1 ^2 ^3Gemini CLI, GitHub (Apache-2.0). github.com/...gemini-cli
  23. ^1 ^2 ^3Google, Gemini API pricing. ai.google.dev/...pricing
  24. ^Aider LLM leaderboards. aider.chat/...leaderboards
  25. ^1 ^2 ^3Aider, GitHub (Apache-2.0). github.com/...aider
  26. ^1 ^2 ^3 ^4Cline, GitHub (Apache-2.0). github.com/...cline
  27. ^1 ^2 ^3Augment Code pricing. augmentcode.com/pricing
  28. ^1 ^2Augment Code, #1 open-source agent on SWE-Bench Verified by combining Claude 3.7 and O1 (31 March 2025). augmentcode.com/...-by-combining-claude-3-7-and-o1
  29. ^Artificial Analysis, Terminal-Bench v2.1 (Wayback Machine capture of 16 July 2026). web.archive.org/...terminalbench-v2-1
  30. ^Google Cloud, Gemini Code Assist Standard and Enterprise overview (Antigravity migration note). docs.cloud.google.com/...overview
  31. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17Terminal-Bench 4.0 leaderboard, tbench.ai (accessed 23 September 2026). tbench.ai/...4.0
  32. ^Ryan Marten, Terminal-Bench 4.0, Terminal-Bench blog (28 August 2026). tbench.ai/...terminal-bench-4-0
  33. ^Terminal-Bench 2.1, Terminal-Bench blog (6 May 2026). tbench.ai/...terminal-bench-2-1
  34. ^Ryan Marten, Alex Shaw, and Andy Konwinski, Terminal-Bench 3.0 measures agent abilities at the frontier, Terminal-Bench blog (30 July 2026). tbench.ai/...terminal-bench-3-0
  35. ^1 ^2 ^3 ^4 ^5Anthropic, Introducing Claude Opus 5.5 (22 September 2026). anthropic.com/claude-opus-5-5
  36. ^Anthropic, Introducing Claude Fable 5.1 and Claude Mythos 5.1 (September 2026). anthropic.com/claude-fable-and-mythos-5-1
  37. ^Anthropic, Introducing Claude Opus 5 (24 July 2026). anthropic.com/...claude-opus-5
  38. ^1 ^2OpenAI, GPT-6 Astra: A new generation of intelligence (3 September 2026). openai.com/...gpt-6-astra
  39. ^1 ^2 ^3OpenAI, Introducing GPT-6 Sol and Luna (22 September 2026). openai.com/...introducing-gpt-6-sol-and-luna
  40. ^1 ^2 ^3OpenAI, API changelog. developers.openai.com/...changelog
  41. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14Artificial Analysis, Terminal-Bench 4.0 Benchmark Leaderboard (accessed 23 September 2026). artificialanalysis.ai/...terminalbench-v4-0
  42. ^1 ^2 ^3 ^4Dmitry Lyalin and Taylor Mullen, An important update: Transitioning Gemini CLI to Antigravity CLI, Google Developers Blog (19 May 2026). developers.googleblog.com/...li-to-antigravity-cli
  43. ^1 ^2Google, Gemini Code Assist consumer accounts (deprecation notice). developers.google.com/...code-assist-individuals
  44. ^1 ^2 ^3Cursor Docs, Models and Pricing. cursor.com/...models
  45. ^1 ^2Cursor, Cursor is now a part of SpaceX (14 August 2026). cursor.com/...joining-spacex
  46. ^1 ^2 ^3 ^4 ^5 ^6SpaceXAI, Introducing Grok 4.7 (21 September 2026). x.ai/...grok-4-7 (launch-day capture with the original 38.0% figure: web.archive.org/...grok-4-7)
  47. ^Cursor, Introducing Grok 4.6 (August 2026). cursor.com/...grok-4-6
  48. ^1 ^2Cursor Docs, VS Code Migration. cursor.com/...vscode
  49. ^GitHub Changelog, Grok 4.7 is now available in GitHub Copilot (21 September 2026). github.blog/...-is-now-available-in-github-copilot
  50. ^1 ^2Scott Wu and Jeff Wang, Introducing Devin Desktop, Cognition (2 June 2026). cognition.com/...introducing-devin-desktop
  51. ^1 ^2 ^3 ^4 ^5Cognition, Introducing SWE-2: Pushing the Pareto Frontier (10 September 2026). cognition.com/...swe-2
  52. ^Anthropic, Claude Code GitHub repository. github.com/...claude-code
  53. ^Amp, Coding agent and dev environment built for the frontier. ampcode.com
  54. ^1 ^2Anthropic, Claude Code overview. code.claude.com/...overview
  55. ^1 ^2Anthropic, Claude Code license notice ("All rights reserved"). github.com/...LICENSE.md
  56. ^Terminal-Bench 2.1 leaderboard, tbench.ai. Wayback Machine capture of 9 July 2026. web.archive.org/...2.1
  57. ^DeepSeek, DeepSeek-V4-Pro-0813, Hugging Face (MIT license). huggingface.co/...DeepSeek-V4-Pro-0813
  58. ^Zhipu AI (Z.ai), GLM-5.3, Hugging Face (GLM-5.3 license). huggingface.co/...GLM-5.3
  59. ^Moonshot AI, Kimi-K3, Hugging Face (Kimi K3 license). huggingface.co/...Kimi-K3
  60. ^Qwen, Qwen3.8-27B, Hugging Face (Apache-2.0). huggingface.co/...Qwen3.8-27B
  61. ^1 ^2Augment Code, Auggie tops SWE-Bench Pro (4 February 2026). augmentcode.com/...auggie-tops-swe-bench-pro
  62. ^1 ^2Anthropic, What is the Max plan? Claude Help Center. support.claude.com/...11049741-what-is-the-max-plan
  63. ^Terminal-Bench 2.1 leaderboard, tbench.ai. Wayback Machine capture of 18 July 2026. web.archive.org/...2.1
  64. ^Cursor, Introducing Grok 4.7 (21 September 2026). cursor.com/...grok-4-7

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

5 revisions · v6 · 4,834 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: xg07 independent adversarial verification 2026-09-23 (V8); writer audit + 1 material + 12 minor fixed; Grok 4.7 trial-count wording re-checked against tbench.ai data

Cite this page: AI Wiki. "Best AI Coding Assistants." aiwiki.ai, updated 23 Sept 2026, fact-checked 23 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/best_ai_coding_assistants

Suggest edit