AI coding agent
An AI coding agent is an autonomous artificial intelligence system that can independently plan, write, test, debug, and deploy software code with minimal human oversight.
Explore AI Code Generation through related topics and the articles other pages reference most.
Ranked by links from other AI Wiki pages.
Articles that also belong to these categories. Counts cover all of AI Code Generation.
Showing 1-60 of 125 articles
An AI coding agent is an autonomous artificial intelligence system that can independently plan, write, test, debug, and deploy software code with minimal human oversight.
AlphaCode is an artificial intelligence system developed by Google DeepMind that generates computer programs capable of solving competitive programming problems at a human-competitive level
AlphaCode 2 is a competitive-programming system built by Google DeepMind that uses a fine-tuned version of the Gemini family of language models to generate, filter, and rank candidate solutions to algorithmic…
Amazon Q is a family of generative AI-powered assistants from Amazon Web Services (AWS), announced on November 28, 2023, at the AWS re:Invent conference and made generally available on April 30, 2024.
Anima (stylized as AnimaApp and operating at animaapp.com) is an Israeli design-to-code platform that converts user-interface designs into developer-ready frontend code.
Antigravity is an agentic software development platform from Google, launched in public preview on 18 November 2025 alongside the release of gemini 3 pro.
Anysphere is an American artificial intelligence company, headquartered in San Francisco, that develops Cursor (code editor), an AI-native code editor built on a fork of Visual Studio Code.
Augment Code is an enterprise AI coding platform, founded in 2022 by Igor Ostrovsky and Guy Gur-Ari and led by CEO Scott Dietzen, that builds AI agents purpose-built for large, complex codebases.
Autonomous coding refers to the use of artificial intelligence systems that can independently write, debug, test, and maintain software with minimal human intervention.
As of July 2026, the best all-around framework for building agentic LLM applications is LangGraph 1.0, the graph-based orchestration layer that reached its first stable release on October 22, 2025 and runs…
As of 23 September 2026, the two tools at the top of the official Terminal-Bench 4.0 leaderboard are OpenAI's Codex and Anthropic's Claude Code.
BigCodeBench is a Python code generation benchmark of 1,140 function-level programming tasks that require composing 723 distinct function calls from 139 libraries across seven domains
Bolt.new is an AI-powered full-stack web development platform built by StackBlitz that turns a natural language prompt into a complete, running web application inside the browser
Builder.io is a San Francisco visual development platform that pairs a headless, drag-and-drop content management system with a suite of AI design-to-code tools.
CRMArena is an AI benchmark for evaluating large language model agents on professional customer relationship management (CRM) tasks inside a realistic, schema-faithful Salesforce environment.
CRUXEval (Code Reasoning, Understanding, and eXecution Evaluation) is a benchmark designed to measure how well large language models can reason about, understand, and mentally execute short Python programs.
ChatDev is an open-source multi-agent software development framework in which large language model (LLM) agents play role-specialized parts (Chief Executive Officer, Chief Technology Officer, programmer…
Claude Code is an agentic software development product made by Anthropic. It uses models from the Claude family to inspect software projects, propose and apply changes, run development tools, and report…
Claude Code hooks connect Claude Code session events to automation, such as checking proposed tool calls or formatting edits.
Claude Code permissions determine whether a tool may run, whether it requires approval, and whether the application blocks it.
Claude Code Review is a multi-agent code review system developed by Anthropic that automatically analyzes GitHub pull requests for bugs, security vulnerabilities, and logic errors.
Claude Sonnet 4.5 is a multimodal large language model (LLM) developed by Anthropic and released on September 29, 2025, which Anthropic described at launch as "the best coding model in the world." It is a…
Cline is an open-source autonomous coding agent that runs as an extension inside VS Code and several other editors, with Cline reporting more than 11 million installs across the VS Code Marketplace and the…
Code Llama is a family of open-weight large language models specialized for code generation and understanding, released by Meta AI on August 24, 2023.
CodeContests is a competitive programming dataset created by Google DeepMind for training and evaluating machine learning models on algorithmic problem-solving tasks.
CodeGeeX is an open series of multilingual code generation models developed by the Knowledge Engineering Group (KEG) and Data Mining lab at Tsinghua University together with Zhipu AI.
CodeGemma is a family of open code-generation models that Google released in April 2024, built on the first generation of its lightweight Gemma models.
CodeMidas is an agentic data pipeline from Xiaomi's LLM Core team that turns functionality already implemented in open-source codebases into executable reinforcement learning environments for coding agents.
Codeium was an artificial intelligence company that built free, unlimited AI code completion and the Windsurf Editor, the IDE its founders called "the first agentic IDE," before becoming the center of a…
Codestral is a family of code-specialized large language models developed by Mistral AI, beginning with Codestral 22B, released on May 29, 2024
A coding agent is a software system that uses a large language model to carry out programming work by acting on a development environment rather than only producing text for a person to copy.
Cognition AI is an American artificial intelligence company headquartered in San Francisco, California, that builds autonomous AI software engineering agents.
Command Code is a commercial AI coding agent sold by subscription, distributed as a terminal program on the npm registry under the package name command-code and, since September 2026
Continue is an open-source AI code assistant that integrates directly into code editors, letting developers connect any large language model (LLM) and customize AI-powered coding features including…
Cursor is a proprietary AI-assisted code editor and software-development platform made by Anysphere.
Cursor Composer 2.5 is a proprietary agentic coding model built by Anysphere, the company behind the Cursor code editor.
CursorBench is an agentic coding benchmark built and run by Cursor, the AI code editor made by Anysphere.
DeepSWE is a benchmark for AI coding agents made by Datacurve, a San Francisco data company.
DeepSeek Harness is an open-source AI agent harness developed by DeepSeek. Also called dsh, it supplies the software around a language model: the agent loop, tools, sessions, filesystems, permission controls…
DeepSeek V4-Flash is the smaller of the two large language models in the DeepSeek V4 family, a 284-billion-parameter Mixture of Experts model with 13 billion active parameters and a one-million-token context…
DeepSeek-Coder is a family of open-weight code large language models built for code generation, completion, and infilling, developed by the Chinese AI lab DeepSeek (DeepSeek-AI).
Ellipsis is a cloud service for configuring and running large language model agents for software-development work.
Factory is an American AI company that builds autonomous software engineering agents for enterprise engineering teams.
FeatureBench is an execution-based benchmark for measuring how well LLM-powered coding agents handle complex, feature-oriented software development rather than bug fixing.
Fill-in-the-middle (FIM) is a training objective and inference technique that lets an autoregressive language model generate text for a gap in the middle of a document, conditioned on both the text before the…
Freebuff is a free, advertising-funded AI coding agent published by Freebuff, Inc., the San Francisco company that also sells the paid Codebuff coding assistant.
FrontierCode is an agentic coding benchmark created by Cognition, the company behind the Devin coding agent.
GLM-5.3 is a text-input and text-output large language model from Z.ai, available through hosted services and downloadable checkpoints.
GPT-5.1-Codex-Max is a frontier agentic coding model from OpenAI, released on November 19, 2025 for Codex, OpenAI's software engineering agent.
GSO (Global Software Optimization), also styled as a benchmark of "Challenging Software Optimization Tasks for Evaluating SWE-Agents," is an AI benchmark that measures whether AI agents and language models can…
Galileo AI (at the domain usegalileo.ai) was a generative AI design tool that turned plain-text prompts into editable, high-fidelity user interface designs, a workflow it popularized as "text to UI." It was…
Gemini CLI is an open-source, terminal-based AI agent developed by Google that brings Gemini models directly into the command line for coding, file editing, shell automation, and research grounded in real-time…
Gemini Code Assist is an AI-powered coding assistant developed by Google that offers code completion, code generation, and natural-language chat inside integrated development environments (IDEs) and across…
The Gemma 4 Developer Agent Competition is a 2026 Kaggle challenge for AI coding agents, accompanied by a research-paper track. Kaggle lists Google DeepMind as host, while the rules name Google LLC as sponsor.
GitHub Copilot is a hosted artificial intelligence coding assistant developed by GitHub.
GitHub Copilot Workspace was a task-centric, AI-powered developer environment built by GitHub Next, the research and incubation arm of GitHub.
Grok Build is an agentic coding tool and command-line interface (CLI) developed by xAI, the artificial intelligence company founded by Elon Musk.
Grok Code Fast is a family of coding-specialized large language models from xai, the artificial intelligence company founded by elon musk.
HumanEval is a benchmark for measuring whether a code-generating large language model can complete short Python functions so that they pass unit tests. It contains 164 hand-written tasks.
Jules is an autonomous AI coding agent developed by Google under Google Labs, designed to perform software development tasks asynchronously inside secure cloud virtual machines.