53 reviewed model records

AI model comparison finder

Find models that meet your actual constraints, then compare their context, capabilities, pricing, distribution, and lifecycle side by side. Every model card links back to a first-party source.

No sponsored rankingsNo login or API keySources checked July 23, 2026

Finder

Narrow the catalog

Results use reviewed first-party model facts, not benchmark rankings.

Required capabilities (models must offer every checked one)

Side by side

Compare selected models

3 of 4 selected

Side-by-side comparison of GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.6 Flash
Attribute

OpenAI

GPT-5.6 Terra

Anthropic

Claude Sonnet 5

Google

Gemini 3.6 Flash

Lifecycle
Active

No retirement announced

Active

Available until at least Jun 30, 2027

Active

No retirement announced

API / checkpoint IDgpt-5.6-terraclaude-sonnet-5gemini-3.6-flash
Release dateJul 9, 2026Jun 30, 2026Jul 21, 2026
Knowledge cutoffFebruary 16, 2026January 2026Not published
DistributionHosted APIHosted APIHosted API
Context window1.05M tokensLargest1M tokens1.05M tokens
Maximum output128K tokens128K tokens65.5K tokens
Modalitiestext, image to texttext, image, pdf to texttext, image, video, audio, pdf to text
Capabilities
reasoningfunction callingstructured outputsstreamingtool use
adaptive thinkingvisiontool useprompt cachingbatch processing
thinkingfunction callingstructured outputscode executionsearch groundingcontext caching
Standard text price

$2.50 input / $15 output

USD per 1M tokens

Higher rates above 272K input tokens ($5 input / $22.50 output).

Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.

$2 input / $10 output

USD per 1M tokens

$2/$10 through Aug 31, 2026, then $3/$15

Introductory pricing through August 31, 2026; Anthropic lists $3 input / $15 output afterward.

$1.50 input / $7.50 outputLowest input

USD per 1M tokens

Paid-tier text/image/video input price; feature charges are separate.

Open-weight detailsHosted API onlyHosted API onlyHosted API only
CaveatsNone notedNone notedNone noted
Official sourceGPT-5.6 Terra model card Claude models overview Gemini 3.6 Flash model card

Matching models

29 of 53 reviewed records

Anthropic

Claude Fable 5

claude-fable-5
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$10 / $50
Released
Jun 9, 2026
Knowledge cutoff
January 2026

Base Claude API price; regional and partner-platform charges can differ.

US export controls forced a global suspension of Fable 5 and Mythos 5 from June 12 to June 30, 2026; access was restored on July 1, 2026. Availability and data-retention eligibility can vary.

Available until at least Jun 9, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001
Active
Context
200K
Max output
64K
Distribution
Hosted API
Input / output price
$1 / $5
Released
Oct 15, 2025
Knowledge cutoff
February 2025

Base Claude API price; regional and partner-platform charges can differ.

The dated API ID is pinned. The shorter claude-haiku-4-5 value is a convenience alias.

Available until at least Oct 15, 2026

text inputimage inputpdf inputextended thinkingvisiontool use
Official source

Anthropic

Claude Mythos 5

claude-mythos-5
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$10 / $50
Released
Jun 9, 2026
Knowledge cutoff
January 2026

Anthropic lists the same specs and pricing as Claude Fable 5.

Not generally available: offered in limited availability to approved Project Glasswing customers for defensive cybersecurity work. Also covered by the June 12 to June 30, 2026 export-control suspension; access was restored July 1, 2026.

No retirement announced

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Opus 4.8

claude-opus-4-8
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$5 / $25
Released
May 28, 2026
Knowledge cutoff
January 2026

Base Claude API price; regional and partner-platform charges can differ.

Available until at least May 28, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use
Official source

Anthropic

Claude Sonnet 5

claude-sonnet-5
Active
Context
1M
Max output
128K
Distribution
Hosted API
Input / output price
$2 / $10
Released
Jun 30, 2026
Knowledge cutoff
January 2026

$2/$10 through Aug 31, 2026, then $3/$15

Introductory pricing through August 31, 2026; Anthropic lists $3 input / $15 output afterward.

Available until at least Jun 30, 2027

text inputimage inputpdf inputadaptive thinkingvisiontool use

DeepSeek

DeepSeek-V4-Flash

deepseek-v4-flash
Active
Context
1M
Max output
384K
Distribution
API and downloadable weights
Input / output price
$0.14 / $0.28
Released
Apr 24, 2026
Knowledge cutoff
Not published

Cache-miss input price; cache hits bill at $0.0028 per 1M. One price covers thinking and non-thinking modes.

384K is the published maximum output; DeepSeek does not list a default output limit, knowledge cutoff, or deprecation commitments.

No retirement announced

text inputreasoningself-hosting
Official source

DeepSeek

DeepSeek-V4-Pro

deepseek-v4-pro
Active
Context
1M
Max output
384K
Distribution
API and downloadable weights
Input / output price
$0.43 / $0.87
Released
Apr 24, 2026
Knowledge cutoff
Not published

Cache-miss input price; cache hits bill at $0.003625 per 1M. One price covers thinking and non-thinking modes.

384K is the published maximum output; DeepSeek does not list a default output limit, knowledge cutoff, or deprecation commitments.

No retirement announced

text inputreasoningself-hosting
Official source

Google

Gemini 3.5 Flash

gemini-3.5-flash
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$1.50 / $9
Released
May 19, 2026
Knowledge cutoff
Not published

Paid-tier text/image/video input price; feature charges are separate.

gemini-flash-latest is a rolling alias and can change target. Use the stable model ID for production pinning.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Official source

Google

Gemini 3.5 Flash-Lite

gemini-3.5-flash-lite
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$0.30 / $2.50
Released
Jul 21, 2026
Knowledge cutoff
Not published

Paid-tier text/image/video input price; feature charges are separate. Context caching is paid-tier only (plus $1.00 per 1M tokens per hour of cache storage).

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Official source

Google

Gemini 3.6 Flash

gemini-3.6-flash
Active
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$1.50 / $7.50
Released
Jul 21, 2026
Knowledge cutoff
Not published

Paid-tier text/image/video input price; feature charges are separate.

No retirement announced

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs

Google

Gemma 4 31B

gemma-4-31b-it
Active
Context
256K
Max output
Not published
Distribution
API and downloadable weights
Input / output price
Varies by host
Released
Apr 2, 2026
Knowledge cutoff
Not published

Open-weight hosting has no universal price or output cap. The first-party Gemini API deployment can have separate service limits.

No retirement announced

text inputimage inputreasoningfunction callingcoding
Official source

OpenAI

GPT-5.4

gpt-5.4
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$2.50 / $15
Released
Not published
Knowledge cutoff
August 31, 2025

Higher rates above 272K input tokens ($5 input / $22.50 output).

Positioned by OpenAI as the affordable coding and professional-work model. No published launch date; the dated snapshot id suggests March 5, 2026.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-5.4-mini

gpt-5.4-mini
Active
Context
400K
Max output
128K
Distribution
Hosted API
Input / output price
$0.75 / $4.50
Released
Not published
Knowledge cutoff
August 31, 2025

No published launch date; the dated snapshot id suggests March 17, 2026.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-5.4-nano

gpt-5.4-nano
Active
Context
400K
Max output
128K
Distribution
Hosted API
Input / output price
$0.20 / $1.25
Released
Not published
Knowledge cutoff
August 31, 2025

No published launch date; the dated snapshot id suggests March 17, 2026.

No retirement announced

text inputimage inputfunction callingstructured outputsstreaming
Official source

OpenAI

GPT-5.5

gpt-5.5
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$5 / $30
Released
Not published
Knowledge cutoff
December 1, 2025

Higher rates above 272K input tokens ($10 input / $45 output).

Still sold and priced, but no longer shown on OpenAI's models index. OpenAI has not published a launch date; the dated snapshot id suggests April 23, 2026.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-5.6 Luna

gpt-5.6-luna
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$1 / $6
Released
Jul 9, 2026
Knowledge cutoff
February 16, 2026

Higher rates above 272K input tokens ($2 input / $9 output).

Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-5.6 Sol

gpt-5.6-sol
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$5 / $30
Released
Jul 9, 2026
Knowledge cutoff
February 16, 2026

Higher rates above 272K input tokens ($10 input / $45 output).

Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.

The gpt-5.6 alias currently resolves to Sol; aliases are mutable, so pin gpt-5.6-sol when reproducibility matters.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

GPT-5.6 Terra

gpt-5.6-terra
Active
Context
1.05M
Max output
128K
Distribution
Hosted API
Input / output price
$2.50 / $15
Released
Jul 9, 2026
Knowledge cutoff
February 16, 2026

Higher rates above 272K input tokens ($5 input / $22.50 output).

Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs

xAI

Grok 4.3

grok-4.3
Active
Context
1M
Max output
Not published
Distribution
Hosted API
Input / output price
$1.25 / $2.50
Released
Not published
Knowledge cutoff
Not published

Higher rates above 200K input tokens ($2.50 input / $5 output).

grok-latest currently resolves to Grok 4.3, not Grok 4.5. xAI does not publish a release date, max-output limit, or deprecation commitments.

No retirement announced

text inputimage input
Official source

xAI

Grok 4.5

grok-4.5
Active
Context
500K
Max output
Not published
Distribution
Hosted API
Input / output price
$2 / $6
Released
Jul 16, 2026
Knowledge cutoff
February 1, 2026

Higher rates above 200K input tokens ($4 input / $12 output).

xAI does not publish max-output limits or deprecation commitments; capability flags are listed on the official model page rather than restated here.

No retirement announced

text inputimage input
Official source

xAI

Grok Build 0.1

grok-build-0.1
Active
Context
256K
Max output
Not published
Distribution
Hosted API
Input / output price
$1 / $2
Released
Not published
Knowledge cutoff
Not published

Higher rates above 200K input tokens ($2 input / $4 output).

Coding-focused model; legacy grok-code-fast IDs resolve here per xAI's docs. xAI does not publish a release date, max-output limit, or deprecation commitments.

No retirement announced

text inputimage inputcoding
Official source

Mistral AI

Mistral Large 3

mistral-large-2512
Active
Context
256K
Max output
Not published
Distribution
API and downloadable weights
Input / output price
$0.50 / $1.50
Released
Dec 2, 2025
Knowledge cutoff
Not published

First-party Mistral API price. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.

mistral-large-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

Mistral AI

Mistral Medium 3.5

mistral-medium-3-5
Active
Context
256K
Max output
Not published
Distribution
API and downloadable weights
Input / output price
$1.50 / $7.50
Released
Apr 28, 2026
Knowledge cutoff
Not published

First-party Mistral API price. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.

Mistral publishes dateless IDs only (mistral-medium-3-5 plus the mistral-medium-3 and mistral-medium-latest aliases); the release date and license wording come from the model card alone.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

Mistral AI

Mistral Small 4

mistral-small-2603
Active
Context
256K
Max output
Not published
Distribution
API and downloadable weights
Input / output price
$0.15 / $0.60
Released
Mar 16, 2026
Knowledge cutoff
Not published

First-party Mistral API price; self-hosting cost is deployment-specific. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.

mistral-small-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.

No retirement announced

text inputimage inputreasoningfunction callingstructured outputs
Official source

OpenAI

gpt-oss-120b

gpt-oss-120b
Active
Context
131.1K
Max output
131.1K
Distribution
Open weights
Input / output price
Varies by host
Released
Aug 5, 2025
Knowledge cutoff
Not published

Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

No retirement announced

text inputreasoningfunction callingstructured outputs
Official source

OpenAI

gpt-oss-20b

gpt-oss-20b
Active
Context
131.1K
Max output
131.1K
Distribution
Open weights
Input / output price
Varies by host
Released
Aug 5, 2025
Knowledge cutoff
Not published

Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

No retirement announced

text inputreasoningfunction callingstructured outputs
Official source

Meta

Llama 4 Maverick

llama-4-maverick
Active
Context
1M
Max output
Not published
Distribution
Open weights
Input / output price
Varies by host
Released
Apr 5, 2025
Knowledge cutoff
Not published

Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers can impose different limits.

No retirement announced

text inputimage inputmixture of expertsmultilingualself-hosting
Official source

Meta

Llama 4 Scout

llama-4-scout
Active
Context
10M
Max output
Not published
Distribution
Open weights
Input / output price
Varies by host
Released
Apr 5, 2025
Knowledge cutoff
Not published

Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers may expose a smaller context window.

No retirement announced

text inputimage inputmixture of expertsmultilingualself-hosting
Official source

Google

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview
Preview
Context
1.05M
Max output
65.5K
Distribution
Hosted API
Input / output price
$2 / $12
Released
Feb 19, 2026
Knowledge cutoff
Not published

Higher rates above 200K input tokens ($4 input / $18 output).

Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Preview models can change or retire on shorter notice. gemini-pro-latest is a rolling alias, not a pinned release.

No retirement date published

text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Official source

Specifications narrow a shortlist; they do not rank quality.

Benchmarks, latency, rate limits, regional availability, safety behavior, and real workload quality still require testing. For workload-specific token costs, use the API cost and context planner.

Method and limits

Facts are versioned; aliases and prices can move

The catalog is a curated snapshot verified on July 23, 2026. It preserves model IDs, effective dates, caveats, and source links rather than silently treating a rolling alias as a permanent checkpoint.

Prices shown are standard base text-token rates where the provider publishes one. Long-context tiers, regions, caching, batches, media, tools, fine-tuning, and hosted open-model prices may differ. Confirm the linked source before committing a budget or migration.

Editorial comparisons

The finder compares published specifications. These wiki articles add the qualitative context: strengths, pricing history, and deployment tradeoffs.

Questions about comparing AI models

What does the AI model comparison include?

The finder compares first-party published facts: lifecycle state, model or deployment ID, context window, maximum output, input and output modalities, selected capabilities, standard text-token pricing, distribution, parameter counts, and licenses where applicable.

Does the cheapest model always cost less for a real workload?

No. Token counts, prompt caching, batch discounts, long-context tiers, tool calls, images, audio, reasoning tokens, retries, and provider-specific fees can change the bill. Use the linked cost planner with a realistic workload.

Does open-weight mean open-source?

Not necessarily. Open-weight means downloadable model weights are available. Each model still has its own license, and some community licenses impose conditions that differ from standard open-source licenses.

Which model is best?

There is no universal winner. This tool narrows models by documented constraints; it does not invent one composite quality score. Test the finalists on representative prompts and measure quality, latency, reliability, and total cost.