73 reviewed model records

AI model comparison finder

Find models that meet your actual constraints, then compare their context, capabilities, pricing, distribution, and lifecycle side by side. Every model card links back to a first-party source.

No sponsored rankingsNo login or API keyLatest fact review October 8, 2026

73 comparison records · Browse all model articles. Article coverage is broader than the reviewed tool catalog. Prices, capabilities, and lifecycle each have their own source and checked date.

Your model workspace

3 selected. Your shortlist and workload carry between tools.

Finder

Narrow the catalog

Results use reviewed first-party model facts, not benchmark rankings.

Required capabilities (models must offer every checked one)

Side by side

Compare selected models

3 of 4 selected

Side-by-side comparison of GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.6 Flash
Attribute

OpenAI

GPT-5.6 Terra

Anthropic

Claude Sonnet 5

Google

Gemini 3.6 Flash

Lifecycle
Active

No retirement announced

Active

Available until at least Jun 30, 2027

Active

No retirement announced

API / checkpoint IDgpt-5.6-terraclaude-sonnet-5gemini-3.6-flash
Release dateJul 9, 2026Jun 30, 2026Jul 21, 2026
Knowledge cutoffFebruary 16, 2026January 2026Not published
DistributionHosted APIHosted APIHosted API
Context window1.05M tokensLargest1M tokens1.05M tokens
Maximum output128K tokens128K tokens65.5K tokens
Modalitiestext, image to texttext, image, pdf to texttext, image, video, audio, pdf to text
Capabilities
reasoningfunction callingstructured outputsstreamingtool use
adaptive thinkingvisiontool useprompt cachingbatch processing
thinkingfunction callingstructured outputscode executionsearch groundingcontext caching
Standard text price

$2 input / $12 output

USD per 1M tokens

Price checked 2026-09-11

Higher rates above 272K input tokens ($4 input / $18 output).

$2 input / $10 output

USD per 1M tokens

Price checked 2026-09-11

These are standard prices. The previously announced September 1 increase was cancelled.

$0.75 input / $3.75 outputLowest input

USD per 1M tokens

Price checked 2026-10-08

$0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50

Google's Gemini 3.7 Flash guide says the 3.7 introductory rate also applies to 3.6 Flash. The first dated listing is the 3.7 Flash model card published August 13, 2026; Google states no separate start date. Launch list price was $1.50 input and $7.50 output. Cache storage and grounding are billed separately.

Open-weight detailsHosted API onlyHosted API onlyHosted API only
CaveatsNone noted

Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends Claude Sonnet 5.5.

The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

Official sourceGPT-5.6 Terra model card Claude sonnet 5 model page Gemini 3.6 Flash model card

Matching models

53 of 73 reviewed records

Anthropic

Claude Fable 5

claude-fable-5
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$10 / $50
Released
Jun 9, 2026
Knowledge cutoff
January 2026

Base Claude API price; regional and partner-platform charges can differ.

Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends Claude Fable 5.1. US export controls forced a global suspension of Fable 5 and Mythos 5 from June 12 to June 30, 2026; access was restored on July 1, 2026.

Available until at least Jun 9, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Fable 5.1

claude-fable-5-1
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$10 / $50
Released
Sep 1, 2026
Knowledge cutoff
June 2026

Base Claude API price. Cache reads cost 0.025x the base input price on this model. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.

Anthropic requires 30-day data retention for this model; it is not available under zero data retention unless Anthropic expressly authorizes it. Forced tool use (tool_choice any or tool) returns a 400 error.

Available until at least Sep 1, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001
Active
Context
200K
Max output
64K
Distribution
Hosted API
Input / output price
$1 / $5
Released
Oct 15, 2025
Knowledge cutoff
February 2025

Base Claude API price; regional and partner-platform charges can differ.

The dated API ID is pinned; claude-haiku-4-5 is a convenience alias. Anthropic's lifecycle table lists it as active with a not-sooner-than date of October 15, 2026, and its model page labels it legacy and recommends Claude Haiku 5.5.

Available until at least Oct 15, 2026

text inputimage inputpdf inputextended thinkingvisiontool use
Official source

Anthropic

Claude Haiku 5.5

claude-haiku-5-5
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$0.10 / $0.50
Released
Oct 7, 2026
Knowledge cutoff
June 2026

Higher rates above 100K input tokens ($0.50 input / $2.50 output).

Base Claude API price for prompts up to 100,000 tokens. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.

Unlike other current Claude models, Haiku 5.5 is priced by prompt length: prompts over 100,000 tokens cost $0.50 input and $2.50 output per million tokens. Manual extended thinking (budget_tokens) returns a 400 error.

Available until at least Oct 7, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Mythos 5

claude-mythos-5
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$10 / $50
Released
Jun 9, 2026
Knowledge cutoff
January 2026

Anthropic lists the same specs and pricing as Claude Fable 5.

Not generally available: Anthropic offers it only to organizations verified through its verification programs, such as the Cyber Verification Program, which Anthropic merged with Project Glasswing into one expanded program on October 6, 2026; Claude Mythos 5.1 is now the current Mythos model. Also covered by the June 12 to June 30, 2026 export-control suspension; access was restored July 1, 2026.

Available until at least Jun 9, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Mythos 5.1

claude-mythos-5-1
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$10 / $50
Released
Sep 1, 2026
Knowledge cutoff
June 2026

Anthropic lists the same specifications and pricing as Claude Fable 5.1, including cache reads at 0.025x the base input price.

Not generally available: Anthropic offers it only to organizations verified through its verification programs, such as the Cyber Verification Program (whose three access tiers, announced October 6, 2026, each include it) and the Life Sciences Verification Program; at launch it was limited to a set of US organizations. It is the same model as Claude Fable 5.1 with different safeguards, and it carries 30-day data retention unless Anthropic expressly authorizes zero data retention.

Available until at least Sep 1, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Opus 4.5

claude-opus-4-5-20251101
Active
Context
200K
Max output
64K
Distribution
Hosted API
Input / output price
$5 / $25
Released
Nov 24, 2025
Knowledge cutoff
May 2025

Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Available until at least Nov 24, 2026

text inputimage inputpdf inputextended thinkingvisiontool use
Official source

Anthropic

Claude Opus 4.6

claude-opus-4-6
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$5 / $25
Released
Feb 5, 2026
Knowledge cutoff
May 2025

Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Available until at least Feb 5, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Opus 4.7

claude-opus-4-7
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$5 / $25
Released
Apr 16, 2026
Knowledge cutoff
January 2026

Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models. Fast mode was deprecated June 25, 2026, with removal on July 24, 2026.

Available until at least Apr 16, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Opus 4.8

claude-opus-4-8
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$5 / $25
Released
May 28, 2026
Knowledge cutoff
January 2026

Base Claude API price; regional and partner-platform charges can differ.

Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends Claude Opus 5.5.

Available until at least May 28, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Opus 5

claude-opus-5
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$5 / $25
Released
Jul 24, 2026
Knowledge cutoff
May 2026

Base Claude API price, the same as Claude Opus 4.8. Fast mode (research preview, Claude API only) costs $10 input and $50 output per million tokens. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.

Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends migrating to Claude Opus 5.5. Disabling thinking is allowed only at effort high or below.

Available until at least Jul 24, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Opus 5.5

claude-opus-5-5
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$4 / $20
Released
Sep 22, 2026
Knowledge cutoff
June 2026

Base Claude API price. Cache reads cost 0.05x the base input price on this model. Fast mode (research preview, Claude API only) costs $8 input and $40 output per million tokens. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.

Adaptive thinking is always on and cannot be turned off, and forced tool use (tool_choice any or tool) returns a 400 error. The Message Batches API allows up to 300K output tokens with the output-300k-2026-03-24 beta header.

Available until at least Sep 22, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Sonnet 4.6

claude-sonnet-4-6
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$3 / $15
Released
Feb 17, 2026
Knowledge cutoff
August 2025

Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Available until at least Feb 17, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Sonnet 5

claude-sonnet-5
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$2 / $10
Released
Jun 30, 2026
Knowledge cutoff
January 2026

These are standard prices. The previously announced September 1 increase was cancelled.

Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends Claude Sonnet 5.5.

Available until at least Jun 30, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use

Anthropic

Claude Sonnet 5.5

claude-sonnet-5-5
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$2 / $10
Released
Sep 28, 2026
Knowledge cutoff
June 2026

Base Claude API price. Cache reads were $0.20 per million tokens at launch on September 28, 2026 and were cut to $0.10 (0.05x the base input price) on October 7, 2026; all other rates are unchanged since launch. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.

Forced tool use (tool_choice any or tool) returns a 400 error, and its thinking blocks work only in the account that produced them or a linked account. The Message Batches API allows up to 300K output tokens with the output-300k-2026-03-24 beta header.

Available until at least Sep 28, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

DeepSeek

DeepSeek-V4-Pro

deepseek-v4-pro
Active
Context
1M
Max output
384K
Distribution
API and downloadable weights
Input / output price
$1.32 / $3.96
Released
Apr 24, 2026
Knowledge cutoff
Not published

Peak rates used for planning. Off-peak input/cache/output: $0.66/$0.022/$1.98 (half the peak rates). Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays; weekends and Chinese public holidays are off-peak all day.

Served by the DeepSeek-V4-Pro-0813 build since August 13, 2026; image input is not supported. DeepSeek's September 10 news post still says this name routes to V4.1-Flash from September 14, but its change log says V4-Pro API service continues after that date with billing unchanged until further notice, and the pricing page still lists it.

No retirement announced

text inputreasoningtool usestructured outputs
Official source

DeepSeek

DeepSeek-V4.1-Flash

deepseek-flash
Active
Context
1M
Max output
384K
Distribution
API and downloadable weights
Input / output price
$0.30 / $1.20
Released
Sep 10, 2026
Knowledge cutoff
Not published

Peak rates used for planning. Off-peak input/cache/output: $0.15/$0.003/$0.60 (half the peak rates). Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays; weekends and Chinese public holidays are off-peak all day.

API model ID is deepseek-flash; the retired names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, served by this model and billed at these rates. Generic dense-transformer memory formulas do not model this architecture accurately.

No retirement announced

text inputimage inputreasoningtool usestructured outputs
Official source

Google

Gemini 2.5 Flash

gemini-2.5-flash
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$0.30 / $2.50
Released
Jun 17, 2025
Knowledge cutoff
Not published

Google's September 18, 2026 release notes limit access to the 2.5 models to users who have actively used them in the past. Google says they are not deprecated and lists no shutdown date, and it directs new projects to Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Official source

Google

Gemini 2.5 Flash-Lite

gemini-2.5-flash-lite
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$0.10 / $0.40
Released
Jul 22, 2025
Knowledge cutoff
Not published

Google's September 18, 2026 release notes limit access to the 2.5 models to users who have actively used them in the past. Google says they are not deprecated and lists no shutdown date, and it directs new projects to Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Official source

Google

Gemini 2.5 Pro

gemini-2.5-pro
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$1.25 / $10
Released
Jun 17, 2025
Knowledge cutoff
Not published

Higher rates above 200K input tokens ($2.50 input / $15 output).

Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Google's September 18, 2026 release notes limit access to the 2.5 models to users who have actively used them in the past. Google says they are not deprecated and lists no shutdown date, and it directs new projects to Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Official source

Google

Gemini 3.5 Flash

gemini-3.5-flash
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$1.50 / $9
Released
May 19, 2026
Knowledge cutoff
Not published

Paid-tier text/image/video input price; feature charges are separate.

Google's models page describes this as its legacy Flash model, but it is still listed as stable and the deprecation table has no shutdown date. Google's May 19, 2026 release notes made it the model behind gemini-flash-latest; latest aliases can be hot-swapped on new releases, so pin the stable ID.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Official source

Google

Gemini 3.5 Flash-Lite

gemini-3.5-flash-lite
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$0.30 / $2.50
Released
Jul 21, 2026
Knowledge cutoff
Not published

Paid-tier text/image/video input price; feature charges are separate. Context caching is paid-tier only (plus $1.00 per 1M tokens per hour of cache storage).

The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Official source

Google

Gemini 3.6 Flash

gemini-3.6-flash
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$0.75 / $3.75
Released
Jul 21, 2026
Knowledge cutoff
Not published

$0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50

Google's Gemini 3.7 Flash guide says the 3.7 introductory rate also applies to 3.6 Flash. The first dated listing is the 3.7 Flash model card published August 13, 2026; Google states no separate start date. Launch list price was $1.50 input and $7.50 output. Cache storage and grounding are billed separately.

The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs

Google

Gemini 3.7 Flash

gemini-3.7-flash
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$0.75 / $3.75
Released
Aug 13, 2026
Knowledge cutoff
March 2026

$0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50

Introductory paid-tier Standard rates through December 31, 2026; the output price includes thinking tokens. Cache storage ($0.50 per 1M tokens per hour) and grounding are billed separately.

Google qualifies the March 2026 knowledge cutoff: some domains may be limited to January 2025. Google says 3.7 Flash remains fully supported for efficiency-first workloads alongside the newer 3.8 Flash.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Official source

Google

Gemini 3.8 Flash

gemini-3.8-flash
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$0.75 / $3.75
Released
Sep 2, 2026
Knowledge cutoff
March 2026

$0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50

Introductory paid-tier Standard rates announced at launch and valid through December 31, 2026; the output price includes thinking tokens. Cache storage ($0.50 per 1M tokens per hour) and grounding are billed separately.

Google qualifies the March 2026 knowledge cutoff: some domains may be limited to January 2025. Google warns the model may use more tokens on complex tasks, especially at higher effort levels.

No retirement announced

text inputimage inputaudio inputvideo inputpdf inputthinkingfunction callingstructured outputs
Official source

Google

Gemma 4 31B

gemma-4-31b-it
Active
Context
256K
Max output
Not published
Distribution
API and downloadable weights
Input / output price
Varies by host
Released
Apr 2, 2026
Knowledge cutoff
Not published

Open-weight hosting has no universal price or output cap. The first-party Gemini API deployment can have separate service limits.

No retirement announced

text inputimage inputreasoningfunction callingcoding
Official source

OpenAI

GPT-4o

gpt-4o
Active
Context
128K
Max output
16.4K
Distribution
Hosted API
Input / output price
$2.50 / $10
Released
May 13, 2024
Knowledge cutoff
October 1, 2023

Rates for the gpt-4o alias as listed on October 8, 2026; the pricing page does not say when they took effect. The deprecated gpt-4o-2024-05-13 snapshot bills $5 input and $15 output until it shuts down.

Only the older gpt-4o-2024-05-13 snapshot is deprecated, with an October 23, 2026 shutdown. The gpt-4o alias points to gpt-4o-2024-08-06 and is not on OpenAI's deprecation list.

No retirement announced

text inputimage inputfunction callingstructured outputsstreaming
Official source

OpenAI

GPT-5.4

gpt-5.4
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$2.50 / $15
Released
Not published
Knowledge cutoff
August 31, 2025

Higher rates above 272K input tokens ($5 input / $22.50 output).

Positioned by OpenAI as the affordable coding and professional-work model. No published launch date; the dated snapshot id suggests March 5, 2026.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-5.4-mini

gpt-5.4-mini
Active
Context
400K
Max output
128K
Distribution
Hosted API
Input / output price
$0.75 / $4.50
Released
Not published
Knowledge cutoff
August 31, 2025

No published launch date; the dated snapshot id suggests March 17, 2026.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-5.5

gpt-5.5
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$5 / $30
Released
Not published
Knowledge cutoff
December 1, 2025

Higher rates above 272K input tokens ($10 input / $45 output).

Still sold and priced, but no longer shown on OpenAI's models index. OpenAI has not published a launch date; the dated snapshot id suggests April 23, 2026.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-5.6 Luna

gpt-5.6-luna
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$0.20 / $1.20
Released
Jul 9, 2026
Knowledge cutoff
February 16, 2026

Higher rates above 272K input tokens ($0.40 input / $1.80 output).

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-5.6 Sol

gpt-5.6-sol
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$4 / $20
Released
Jul 9, 2026
Knowledge cutoff
February 16, 2026

Higher rates above 272K input tokens ($8 input / $30 output).

Promotional rates are available at least through November 21, 2026. Later rates have not been committed; projections hold these rates constant.

The gpt-5.6 alias currently resolves to Sol; aliases are mutable, so pin gpt-5.6-sol when reproducibility matters.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-5.6 Terra

gpt-5.6-terra
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$2 / $12
Released
Jul 9, 2026
Knowledge cutoff
February 16, 2026

Higher rates above 272K input tokens ($4 input / $18 output).

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs

OpenAI

GPT-6 Astra

gpt-6-astra
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$10 / $50
Released
Sep 3, 2026
Knowledge cutoff
April 30, 2026

Higher rates above 272K input tokens ($20 input / $75 output).

Flex is priced like Batch and Fast mode costs 2x Standard. Since September 29, 2026 an Ultrafast tier (service_tier ultrafast, Responses API) costs $60 input, $6 cached input, $75 cache writes and $300 output per 1M tokens, or $120, $12, $150 and $450 above 272K input tokens.

Does not accept the none reasoning effort, custom temperature or top_p values, or log probabilities, and tool calling requires the Responses API (Chat Completions works without tools).

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-6 Luna

gpt-6-luna
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$0.10 / $0.50
Released
Sep 22, 2026
Knowledge cutoff
May 18, 2026

Higher rates above 272K input tokens ($0.20 input / $0.75 output).

Flex is priced like Batch, Fast mode costs 2x Standard, and regional processing adds 10 percent where available.

On September 25, 2026 OpenAI fixed an image-encoding bug that had degraded image understanding, so rerun image evaluations made before then. Chat Completions supports function calling only with reasoning effort none; use the Responses API for tools with reasoning.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-6 Sol

gpt-6-sol
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$2 / $10
Released
Sep 22, 2026
Knowledge cutoff
April 20, 2026

Higher rates above 272K input tokens ($4 input / $15 output).

Flex is priced like Batch, Fast mode costs 2x Standard, and regional processing adds 10 percent where available.

OpenAI's docs now point to GPT-6.1 Sol as the newer Sol model, though gpt-6-sol has no announced deprecation. On September 25, 2026 OpenAI fixed an image-encoding bug that had degraded image understanding, so rerun image evaluations made before then.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-6.1 Sol

gpt-6.1-sol
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$2 / $10
Released
Sep 29, 2026
Knowledge cutoff
April 30, 2026

Higher rates above 272K input tokens ($4 input / $15 output).

Cached input is billed at 5 percent of the input rate, half of GPT-6 Sol's cached price. Flex is priced like Batch, Fast mode costs 2x Standard, and regional processing adds 10 percent where available.

Does not accept the none or minimal reasoning efforts, and tool calling requires the Responses API (Chat Completions works without tools).

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

xAI

Grok 4.3

grok-4.3
Active
Context
1M
Max output
Not published
Distribution
Hosted API
Input / output price
$1.25 / $2.50
Released
Not published
Knowledge cutoff
Not published

Higher rates at or above 200K input tokens ($2.50 input / $5 output).

The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

xAI

Grok 4.5

grok-4.5
Active
Context
500K
Max output
Not published
Distribution
Hosted API
Input / output price
$2 / $6
Released
Jul 16, 2026
Knowledge cutoff
February 1, 2026

Higher rates at or above 200K input tokens ($4 input / $12 output).

The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

xAI

Grok 4.6

grok-4.6
Active
Context
500K
Max output
Not published
Distribution
Hosted API
Input / output price
$2 / $6
Released
Aug 12, 2026
Knowledge cutoff
February 1, 2026

Higher rates at or above 200K input tokens ($4 input / $12 output).

No batch discount: the Batch API is not supported for this model. Requests to the US regional endpoint (https://us.api.x.ai/v1) bill at 1.1x these rates, Priority Processing bills at 2x, and server-side tool calls are billed separately.

xAI states that this model has no text output limit and no Batch API support, and xAI publishes no retirement date for it. Its overview page now redirects to Grok 4.7 and the models page recommends Grok 4.7, but grok-4.6 keeps its model page, its pricing row and its place on the US regional endpoint.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

xAI

Grok 4.7

grok-4.7
Active
Context
500K
Max output
Not published
Distribution
Hosted API
Input / output price
$2 / $6
Released
Sep 21, 2026
Knowledge cutoff
May 2026

Higher rates at or above 200K input tokens ($4 input / $12 output).

No batch discount: the Batch API is not supported for this model. Requests to the US regional endpoint (https://us.api.x.ai/v1) bill at 1.1x these rates, Priority Processing bills at 2x, and server-side tool calls are billed separately.

xAI states that this model has no text output limit and no Batch API support, and xAI publishes no retirement date for it. Grok 4.7 Fast (the same model at 2x standard rates, or 1.5x for long-context requests) is sold only in Cursor and Grok Build, not on the public xAI API.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

xAI

Grok Build 0.1

grok-build-0.1
Active
Context
256K
Max output
Not published
Distribution
Hosted API
Input / output price
$1 / $2
Released
May 29, 2026
Knowledge cutoff
Not published

Higher rates at or above 200K input tokens ($2 input / $4 output).

xAI announced grok-build-0.1 on its API in public beta on May 29, 2026 (its May release notes call it early access). The official model page lists no maximum output limit or committed retirement date, and alias mappings can change, so use the explicit model ID.

No retirement announced

text inputimage inputcodingreasoningfunction calling
Official source

Mistral AI

Mistral Large 3

mistral-large-2512
Active
Context
256K
Max output
Not published
Distribution
API and downloadable weights
Input / output price
$0.50 / $1.50
Released
Dec 2, 2025
Knowledge cutoff
Not published

First-party Mistral API price. Cached-input and batch rates are published directly on Mistral's pricing page (Standard and Batch views).

mistral-large-latest is a rolling alias that still points to Mistral Large 3 (mistral-large-2512), not to the Mistral Large 4 public preview. The published model card does not specify a separate maximum-output limit.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

Mistral AI

Mistral Medium 3.5

mistral-medium-3-5
Active
Context
256K
Max output
Not published
Distribution
API and downloadable weights
Input / output price
$1.50 / $7.50
Released
Apr 28, 2026
Knowledge cutoff
Not published

First-party Mistral API price. Cached-input and batch rates are published directly on Mistral's pricing page (Standard and Batch views).

Mistral publishes dateless IDs only (mistral-medium-3-5 plus the mistral-medium-3 and mistral-medium-latest aliases); the release date and license wording come from the model card alone.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

Mistral AI

Mistral Small 4

mistral-small-2603
Active
Context
256K
Max output
Not published
Distribution
API and downloadable weights
Input / output price
$0.15 / $0.60
Released
Mar 16, 2026
Knowledge cutoff
Not published

First-party Mistral API price; self-hosting cost is deployment-specific. Cached-input and batch rates are published directly on Mistral's pricing page (Standard and Batch views).

mistral-small-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

Meta

Muse Spark 1.3

muse-spark-1.3
Active
Context
1.05M
Max output
Not published
Distribution
Hosted API
Input / output price
$1.25 / $4.25
Released
Sep 2, 2026
Knowledge cutoff
Not published

Standard tier rates as observed on October 8, 2026 (Meta's pricing page does not date them); prompts and completions are not used to train Meta models. No long-context premium; built-in web search grounding adds $2.50 per 1,000 queries. The muse-spark-1.3-contributor variant bills $0.10/$0.002/$0.20 in exchange for training use.

Meta lists audio input but says audio understanding in Muse Spark 1.3 is not fully supported and may be degraded (it points audio work to Muse Spark 1.2 or Muse Voice Transcribe). Meta does not publish a maximum output limit, knowledge cutoff, or parameter count, and the max reasoning level is Standard-tier only.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputreasoningfunction callingstructured outputs
Official source

Meta

Muse Spark 1.3 (Contributor tier)

muse-spark-1.3-contributor
Active
Context
1.05M
Max output
Not published
Distribution
Hosted API
Input / output price
$0.10 / $0.20
Released
Not published
Knowledge cutoff
Not published

Contributor tier rates as observed on October 8, 2026 (Meta's pricing page does not date them): discounted in exchange for permission to use prompts and completions to train future Meta models. The same model on the Standard tier (muse-spark-1.3) costs $1.25/$0.15/$4.25. Built-in web search grounding adds $2.50 per 1,000 queries.

Training-eligible tier of Muse Spark 1.3: Meta may use your prompts and completions to train future Meta models, the max reasoning level is not available, and the rate limit is 100 requests and 3,000,000 tokens per minute. Meta does not date this tier's launch, and it says audio understanding in Muse Spark 1.3 is not fully supported.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputreasoningfunction callingstructured outputs
Official source

OpenAI

gpt-oss-120b

gpt-oss-120b
Active
Context
131.1K
Max output
131.1K
Distribution
Open weights
Input / output price
Varies by host
Released
Aug 5, 2025
Knowledge cutoff
Not published

Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

No retirement announced

text inputreasoningfunction callingstructured outputs
Official source

OpenAI

gpt-oss-20b

gpt-oss-20b
Active
Context
131.1K
Max output
131.1K
Distribution
Open weights
Input / output price
Varies by host
Released
Aug 5, 2025
Knowledge cutoff
Not published

Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

No retirement announced

text inputreasoningfunction callingstructured outputs
Official source

Meta

Llama 4 Maverick

llama-4-maverick
Active
Context
1M
Max output
Not published
Distribution
Open weights
Input / output price
Varies by host
Released
Apr 5, 2025
Knowledge cutoff
Not published

Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers can impose different limits.

No retirement announced

text inputimage inputmixture of expertsmultilingualself-hosting
Official source

Meta

Llama 4 Scout

llama-4-scout
Active
Context
10M
Max output
Not published
Distribution
Open weights
Input / output price
Varies by host
Released
Apr 5, 2025
Knowledge cutoff
Not published

Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers may expose a smaller context window.

No retirement announced

text inputimage inputmixture of expertsmultilingualself-hosting
Official source

Google

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview
Preview
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$2 / $12
Released
Feb 19, 2026
Knowledge cutoff
Not published

Higher rates above 200K input tokens ($4 input / $18 output).

Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Preview models can change or retire on shorter notice. gemini-pro-latest is a rolling alias, not a pinned release.

No retirement date published

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Official source

Mistral AI

Mistral Large 4

mistral-large-4
Preview
Context
1M
Max output
Not published
Distribution
Hosted API
Input / output price
$1.36 / $4.18
Released
Oct 6, 2026
Knowledge cutoff
Not published

Standard (original) Mistral API rates; batch cached input is $0.07. Mistral's changelog announces launch pricing of 50% off for 2 weeks, shown on the pricing page as a temporary sale ($0.68 input, $0.07 cached input, $2.09 output; batch $0.34/$0.035/$1.045). Mistral publishes no exact end date for the sale, so planning uses the standard rates.

Public Preview: Mistral allows silent updates and does not guarantee general availability, and the model card lists the weights and license as coming soon (Mistral says weights ship by the end of October 2026). Mistral publishes no maximum output limit, and mistral-large-latest still points to Mistral Large 3.

No retirement date published

text inputimage inputreasoningfunction callingstructured outputs
Official source

Specifications narrow a shortlist; they do not rank quality.

Benchmarks, latency, rate limits, regional availability, safety behavior, and real workload quality still require testing. For workload-specific token costs, use the API cost and context planner.

Method and limits

Facts are versioned; aliases and prices can move

The catalog is reviewed field by field; the latest review was October 8, 2026. Untouched fields retain their earlier checked dates. It preserves model IDs, effective dates, caveats, and source links rather than silently treating a rolling alias as a permanent checkpoint.

Prices shown are standard base text-token rates where the provider publishes one. Long-context tiers, regions, caching, batches, media, tools, fine-tuning, and hosted open-model prices may differ. Confirm the linked source before committing a budget or migration.

Editorial comparisons

The finder compares published specifications. These wiki articles add the qualitative context: strengths, pricing history, and deployment tradeoffs.

Questions about comparing AI models

What does the AI model comparison include?

The finder compares first-party published facts: lifecycle state, model or deployment ID, context window, maximum output, input and output modalities, selected capabilities, standard text-token pricing, distribution, parameter counts, and licenses where applicable.

Does the cheapest model always cost less for a real workload?

No. Token counts, prompt caching, batch discounts, long-context tiers, tool calls, images, audio, reasoning tokens, retries, and provider-specific fees can change the bill. Use the linked cost planner with a realistic workload.

Does open-weight mean open-source?

Not necessarily. Open-weight means downloadable model weights are available. Each model still has its own license, and some community licenses impose conditions that differ from standard open-source licenses.

Which model is best?

There is no universal winner. This tool narrows models by documented constraints; it does not invent one composite quality score. Test the finalists on representative prompts and measure quality, latency, reliability, and total cost.