API migration intelligence

AI model deprecations and retirement tracker

See what is retiring, how long you have, and which model each provider recommends next. Search 73 curated API and open-weight records with release history, context limits, standard text pricing, and links back to official notices.

7 providersOfficial first-party sourcesLatest fact review October 8, 2026; dates vary by fieldCalendar export
Your model workspace

Select models to compare, estimate costs, and check deployment options.

Model lifecycle overview

Next 90 days

9

Retirement or earliest-shutdown milestones in current view.

Deprecated

12

Models in current view that need an active migration plan.

Active

51

Current API and open-weight records in current view.

Open weights

12

Checkpoints in current view; hosting limits and cost vary.

Migration queue

Upcoming lifecycle deadlines

Explore

Find a model or migration

Provider

Showing 73 of 73 curated records

OpenAIDeprecated

GPT-3.5 Turbo

gpt-3.5-turbo

Aliases: gpt-3.5-turbo-0125, gpt-3.5-turbo-completions

Retirement

Oct 23, 2026

13d 16h remaining (as of October 9)

Recommended replacement

GPT-5.6 Terragpt-5.6-terra

Run representative low-latency and high-volume evaluations on Terra and update legacy Chat Completions assumptions.

Context
16.4K
Max output
4.1K
Input / output
$0.50 / $1.50

Capabilities

function callingstreaming

Input: text

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4

gpt-4

Aliases: gpt-4-0613, gpt-4-0613-completions, gpt-4-completions

Retirement

Oct 23, 2026

13d 16h remaining (as of October 9)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-evaluate prompts and tool schemas on GPT-5.6 Sol, then replace every GPT-4 snapshot and Completions alias before shutdown.

Context
8.2K
Max output
8.2K
Input / output
$30 / $60

Capabilities

function callingstreaming

Input: text

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4 Turbo

gpt-4-turbo

Aliases: gpt-4-turbo-2024-04-09, gpt-4-turbo-completions

Retirement

Oct 23, 2026

13d 16h remaining (as of October 9)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Benchmark long prompts and multimodal inputs on GPT-5.6 Sol, then move both the rolling alias and dated snapshot.

Context
128K
Max output
4.1K
Input / output
$10 / $30

Capabilities

function callingstructured outputsstreaming

Input: text, image

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4o (2024-05-13 snapshot)

gpt-4o-2024-05-13

Retirement

Oct 23, 2026

13d 16h remaining (as of October 9)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-test multimodal prompts, tool schemas, and latency assumptions on GPT-5.6 Sol before the October 23, 2026 shutdown.

Context
Not published
Max output
Not published
Input / output
$5 / $15

Capabilities

function callingstreaming

Input: text, image

Output: text

Pricing note: The snapshot's launch rates, still listed by OpenAI until the shutdown.

Caveat: Only this dated snapshot of GPT-4o is retiring. The gpt-4o alias now points to gpt-4o-2024-08-06 and is not on OpenAI's deprecation list.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o1

o1

Aliases: o1-2024-12-17

Retirement

Oct 23, 2026

13d 16h remaining (as of October 9)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-test reasoning effort, token budgets, and tool behavior on Sol before changing the production model ID.

Context
200K
Max output
100K
Input / output
$15 / $60

Capabilities

reasoningfunction callingstructured outputs

Input: text, image

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o1-pro

o1-pro

Aliases: o1-pro-2025-03-19

Retirement

Oct 23, 2026

13d 16h remaining (as of October 9)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

OpenAI's documented migration is gpt-5.6-sol with reasoning.mode set to pro; re-test reasoning workloads and budgets before cutover.

Context
200K
Max output
100K
Input / output
$150 / $600

Capabilities

reasoningstructured outputs

Input: text, image

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Caveat: Responses API only. OpenAI has not published a launch date; the dated snapshot id suggests March 19, 2025.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o3-mini

o3-mini

Aliases: o3-mini-2025-01-31

Retirement

Oct 23, 2026

13d 16h remaining (as of October 9)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Validate reasoning-intensive workloads on Sol and account for its different price and latency profile.

Context
200K
Max output
100K
Input / output
$1.10 / $4.40

Capabilities

reasoningfunction callingstructured outputs

Input: text

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o4-mini

o4-mini

Aliases: o4-mini-2025-04-16

Retirement

Oct 23, 2026

13d 16h remaining (as of October 9)

Recommended replacement

GPT-5.6 Terragpt-5.6-terra

Test Terra with the same tools and reasoning workloads; check quality, latency, and token use before cutover.

Context
200K
Max output
100K
Input / output
$1.10 / $4.40

Capabilities

reasoningfunction callingstructured outputs

Input: text, image

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicDeprecated

Claude Sonnet 4.5

claude-sonnet-4-5-20250929

Aliases: claude-sonnet-4-5

Retirement

Nov 30, 2026

51d 16h remaining (as of October 9)

Announced Sep 30, 2026

Recommended replacement

Claude Sonnet 5.5claude-sonnet-5-5

Anthropic recommends migrating to Claude Sonnet 5.5 before November 30, 2026; the Sonnet 5.5 migration guide has a section for moving from Sonnet 4.5.

Context
200K
Max output
64K
Input / output
$3 / $15

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Deprecated September 30, 2026. The November 30, 2026 retirement applies to Anthropic-operated platforms (Claude API, Claude Platform on AWS, Microsoft Foundry); Amazon Bedrock and Google Cloud set their own dates.

Read the AI Wiki article
Official sources (5)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-5.4-nano

gpt-5.4-nano

Aliases: gpt-5.4-nano-2026-03-17

Retirement

Apr 1, 2027

5 months remaining (as of October 9)

Announced Oct 1, 2026

Recommended replacement

GPT-6 Lunagpt-6-luna

OpenAI's recommended replacement. Set reasoning effort explicitly when migrating: gpt-5.4-nano defaults to none, while GPT-6 Luna defaults to medium.

Context
400K
Max output
128K
Input / output
$0.20 / $1.25

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Caveat: Deprecated on October 1, 2026; OpenAI will remove it from the API on April 1, 2027. No published launch date; the dated snapshot id suggests March 17, 2026.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleDeprecated

Gemini 3.1 Flash-Lite

gemini-3.1-flash-lite

Earliest shutdown

May 7, 2027

6 months remaining (as of October 9)

Recommended replacement

Gemini 3.5 Flash-Litegemini-3.5-flash-lite

Google names Gemini 3.5 Flash-Lite as the successor; compare quality and latency before moving latency-sensitive workloads.

Context
1.05M
Max output
65.5K
Input / output
$0.25 / $1.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; audio input is billed at $0.50 per 1M. Context caching adds $1.00 per 1M tokens per hour of storage.

Caveat: Google lists May 7, 2027 as the earliest possible shutdown and names Gemini 3.5 Flash-Lite as the replacement.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicDeprecated

Claude Mythos Preview

claude-mythos-preview

Recommended replacement

Claude Mythos 5claude-mythos-5

Anthropic's documented successor is Claude Mythos 5, available only to organizations verified through Anthropic's verification programs, such as the Cyber Verification Program; Claude Mythos 5.1 is now the current Mythos model.

Context
1M
Max output
Not published
Input / output
Varies

Capabilities

Input: text, image

Output: text

Caveat: The current lifecycle documentation marks Mythos Preview deprecated and directs users to Mythos 5. It does not provide a retirement date; do not treat the former July date as a current commitment.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GooglePreview

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview

Aliases: gemini-pro-latest

Context
1.05M
Max output
65.5K
Input / output
$2 / $12

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Caveat: Preview models can change or retire on shorter notice. gemini-pro-latest is a rolling alias, not a pinned release.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

MistralPreview

Mistral Large 4

mistral-large-4

Aliases: mistral-large-4-0

Context
1M
Max output
Not published
Input / output
$1.36 / $4.18

Capabilities

reasoningfunction callingstructured outputsbatch processingmultilingual

Input: text, image

Output: text

Parameters: 1.05T total / 52B active (plus a 1.6B vision encoder)

Pricing note: Standard (original) Mistral API rates; batch cached input is $0.07. Mistral's changelog announces launch pricing of 50% off for 2 weeks, shown on the pricing page as a temporary sale ($0.68 input, $0.07 cached input, $2.09 output; batch $0.34/$0.035/$1.045). Mistral publishes no exact end date for the sale, so planning uses the standard rates.

Caveat: Public Preview: Mistral allows silent updates and does not guarantee general availability, and the model card lists the weights and license as coming soon (Mistral says weights ship by the end of October 2026). Mistral publishes no maximum output limit, and mistral-large-latest still points to Mistral Large 3.

Read the AI Wiki article
Official sources (5)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Fable 5

claude-fable-5

Availability commitment

No retirement sooner than Jun 9, 2027

Context
1M
Max output
128K
Input / output
$10 / $50

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends Claude Fable 5.1. US export controls forced a global suspension of Fable 5 and Mythos 5 from June 12 to June 30, 2026; access was restored on July 1, 2026.

Read the AI Wiki article
Official sources (6)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Fable 5.1

claude-fable-5-1

Availability commitment

No retirement sooner than Sep 1, 2027

Context
1M
Max output
128K
Input / output
$10 / $50

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price. Cache reads cost 0.025x the base input price on this model. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.

Caveat: Anthropic requires 30-day data retention for this model; it is not available under zero data retention unless Anthropic expressly authorizes it. Forced tool use (tool_choice any or tool) returns a 400 error.

Read the AI Wiki article
Official sources (6)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Haiku 4.5

claude-haiku-4-5-20251001

Aliases: claude-haiku-4-5

Availability commitment

No retirement sooner than Oct 15, 2026

Context
200K
Max output
64K
Input / output
$1 / $5

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Caveat: The dated API ID is pinned; claude-haiku-4-5 is a convenience alias. Anthropic's lifecycle table lists it as active with a not-sooner-than date of October 15, 2026, and its model page labels it legacy and recommends Claude Haiku 5.5.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Haiku 5.5

claude-haiku-5-5

Availability commitment

No retirement sooner than Oct 7, 2027

Context
1M
Max output
128K
Input / output
$0.10 / $0.50

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price for prompts up to 100,000 tokens. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.

Caveat: Unlike other current Claude models, Haiku 5.5 is priced by prompt length: prompts over 100,000 tokens cost $0.50 input and $2.50 output per million tokens. Manual extended thinking (budget_tokens) returns a 400 error.

Read the AI Wiki article
Official sources (6)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Mythos 5

claude-mythos-5

Availability commitment

No retirement sooner than Jun 9, 2027

Context
1M
Max output
128K
Input / output
$10 / $50

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Anthropic lists the same specs and pricing as Claude Fable 5.

Caveat: Not generally available: Anthropic offers it only to organizations verified through its verification programs, such as the Cyber Verification Program, which Anthropic merged with Project Glasswing into one expanded program on October 6, 2026; Claude Mythos 5.1 is now the current Mythos model. Also covered by the June 12 to June 30, 2026 export-control suspension; access was restored July 1, 2026.

Read the AI Wiki article
Official sources (6)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Mythos 5.1

claude-mythos-5-1

Availability commitment

No retirement sooner than Sep 1, 2027

Context
1M
Max output
128K
Input / output
$10 / $50

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Anthropic lists the same specifications and pricing as Claude Fable 5.1, including cache reads at 0.025x the base input price.

Caveat: Not generally available: Anthropic offers it only to organizations verified through its verification programs, such as the Cyber Verification Program (whose three access tiers, announced October 6, 2026, each include it) and the Life Sciences Verification Program; at launch it was limited to a set of US organizations. It is the same model as Claude Fable 5.1 with different safeguards, and it carries 30-day data retention unless Anthropic expressly authorizes zero data retention.

Read the AI Wiki article
Official sources (8)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Opus 4.5

claude-opus-4-5-20251101

Aliases: claude-opus-4-5

Availability commitment

No retirement sooner than Nov 24, 2026

Context
200K
Max output
64K
Input / output
$5 / $25

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Opus 4.6

claude-opus-4-6

Availability commitment

No retirement sooner than Feb 5, 2027

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Opus 4.7

claude-opus-4-7

Availability commitment

No retirement sooner than Apr 16, 2027

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models. Fast mode was deprecated June 25, 2026, with removal on July 24, 2026.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Opus 4.8

claude-opus-4-8

Availability commitment

No retirement sooner than May 28, 2027

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends Claude Opus 5.5.

Read the AI Wiki article
Official sources (5)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Opus 5

claude-opus-5

Availability commitment

No retirement sooner than Jul 24, 2027

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price, the same as Claude Opus 4.8. Fast mode (research preview, Claude API only) costs $10 input and $50 output per million tokens. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends migrating to Claude Opus 5.5. Disabling thinking is allowed only at effort high or below.

Read the AI Wiki article
Official sources (6)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Opus 5.5

claude-opus-5-5

Availability commitment

No retirement sooner than Sep 22, 2027

Context
1M
Max output
128K
Input / output
$4 / $20

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price. Cache reads cost 0.05x the base input price on this model. Fast mode (research preview, Claude API only) costs $8 input and $40 output per million tokens. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.

Caveat: Adaptive thinking is always on and cannot be turned off, and forced tool use (tool_choice any or tool) returns a 400 error. The Message Batches API allows up to 300K output tokens with the output-300k-2026-03-24 beta header.

Read the AI Wiki article
Official sources (6)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Sonnet 4.6

claude-sonnet-4-6

Availability commitment

No retirement sooner than Feb 17, 2027

Context
1M
Max output
128K
Input / output
$3 / $15

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Sonnet 5

claude-sonnet-5

Availability commitment

No retirement sooner than Jun 30, 2027

Context
1M
Max output
128K
Input / output
$2 / $10

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: These are standard prices. The previously announced September 1 increase was cancelled.

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends Claude Sonnet 5.5.

Read the AI Wiki article
Official sources (5)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Sonnet 5.5

claude-sonnet-5-5

Availability commitment

No retirement sooner than Sep 28, 2027

Context
1M
Max output
128K
Input / output
$2 / $10

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price. Cache reads were $0.20 per million tokens at launch on September 28, 2026 and were cut to $0.10 (0.05x the base input price) on October 7, 2026; all other rates are unchanged since launch. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.

Caveat: Forced tool use (tool_choice any or tool) returns a 400 error, and its thinking blocks work only in the account that produced them or a linked account. The Message Batches API allows up to 300K output tokens with the output-300k-2026-03-24 beta header.

Read the AI Wiki article
Official sources (7)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

DeepSeekActiveAPI + open weights

DeepSeek-V4-Pro

deepseek-v4-pro
Context
1M
Max output
384K
Input / output
$1.32 / $3.96

Capabilities

reasoningtool usestructured outputsself-hosting

Input: text

Output: text

Parameters: 1.6T total / 49B active

License: MIT

Pricing note: Peak rates used for planning. Off-peak input/cache/output: $0.66/$0.022/$1.98 (half the peak rates). Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays; weekends and Chinese public holidays are off-peak all day.

Caveat: Served by the DeepSeek-V4-Pro-0813 build since August 13, 2026; image input is not supported. DeepSeek's September 10 news post still says this name routes to V4.1-Flash from September 14, but its change log says V4-Pro API service continues after that date with billing unchanged until further notice, and the pricing page still lists it.

Read the AI Wiki article
Official sources (10)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

DeepSeekActiveAPI + open weights

DeepSeek-V4.1-Flash

deepseek-flash
Context
1M
Max output
384K
Input / output
$0.30 / $1.20

Capabilities

reasoningtool usestructured outputsself-hosting

Input: text, image

Output: text

Parameters: 552B backbone + 196B Engram; 8B active prefill / 16B decode

License: MIT

Pricing note: Peak rates used for planning. Off-peak input/cache/output: $0.15/$0.003/$0.60 (half the peak rates). Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays; weekends and Chinese public holidays are off-peak all day.

Caveat: API model ID is deepseek-flash; the retired names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, served by this model and billed at these rates. Generic dense-transformer memory formulas do not model this architecture accurately.

Read the AI Wiki article
Official sources (4)
GoogleActive

Gemini 2.5 Flash

gemini-2.5-flash
Context
1.05M
Max output
65.5K
Input / output
$0.30 / $2.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Caveat: Google's September 18, 2026 release notes limit access to the 2.5 models to users who have actively used them in the past. Google says they are not deprecated and lists no shutdown date, and it directs new projects to Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 2.5 Flash-Lite

gemini-2.5-flash-lite
Context
1.05M
Max output
65.5K
Input / output
$0.10 / $0.40

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Caveat: Google's September 18, 2026 release notes limit access to the 2.5 models to users who have actively used them in the past. Google says they are not deprecated and lists no shutdown date, and it directs new projects to Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 2.5 Pro

gemini-2.5-pro
Context
1.05M
Max output
65.5K
Input / output
$1.25 / $10

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Caveat: Google's September 18, 2026 release notes limit access to the 2.5 models to users who have actively used them in the past. Google says they are not deprecated and lists no shutdown date, and it directs new projects to Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.5 Flash

gemini-3.5-flash

Aliases: gemini-flash-latest

Context
1.05M
Max output
65.5K
Input / output
$1.50 / $9

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; feature charges are separate.

Caveat: Google's models page describes this as its legacy Flash model, but it is still listed as stable and the deprecation table has no shutdown date. Google's May 19, 2026 release notes made it the model behind gemini-flash-latest; latest aliases can be hot-swapped on new releases, so pin the stable ID.

Read the AI Wiki article
Official sources (5)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.5 Flash-Lite

gemini-3.5-flash-lite
Context
1.05M
Max output
65.5K
Input / output
$0.30 / $2.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; feature charges are separate. Context caching is paid-tier only (plus $1.00 per 1M tokens per hour of cache storage).

Caveat: The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.6 Flash

gemini-3.6-flash
Context
1.05M
Max output
65.5K
Input / output
$0.75 / $3.75

Current pricing through Dec 31, 2026; $1.50/$7.50 from Jan 1, 2027.

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Google's Gemini 3.7 Flash guide says the 3.7 introductory rate also applies to 3.6 Flash. The first dated listing is the 3.7 Flash model card published August 13, 2026; Google states no separate start date. Launch list price was $1.50 input and $7.50 output. Cache storage and grounding are billed separately.

Caveat: The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

Read the AI Wiki article
Official sources (7)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.7 Flash

gemini-3.7-flash
Context
1.05M
Max output
65.5K
Input / output
$0.75 / $3.75

Current pricing through Dec 31, 2026; $1.50/$7.50 from Jan 1, 2027.

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext cachingbatch processing

Input: text, image, video, audio, pdf

Output: text

Pricing note: Introductory paid-tier Standard rates through December 31, 2026; the output price includes thinking tokens. Cache storage ($0.50 per 1M tokens per hour) and grounding are billed separately.

Caveat: Google qualifies the March 2026 knowledge cutoff: some domains may be limited to January 2025. Google says 3.7 Flash remains fully supported for efficiency-first workloads alongside the newer 3.8 Flash.

Read the AI Wiki article
Official sources (6)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.8 Flash

gemini-3.8-flash
Context
1.05M
Max output
65.5K
Input / output
$0.75 / $3.75

Current pricing through Dec 31, 2026; $1.50/$7.50 from Jan 1, 2027.

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext cachingbatch processing

Input: text, image, audio, video, pdf

Output: text

Pricing note: Introductory paid-tier Standard rates announced at launch and valid through December 31, 2026; the output price includes thinking tokens. Cache storage ($0.50 per 1M tokens per hour) and grounding are billed separately.

Caveat: Google qualifies the March 2026 knowledge cutoff: some domains may be limited to January 2025. Google warns the model may use more tokens on complex tasks, especially at higher effort levels.

Read the AI Wiki article
Official sources (6)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActiveAPI + open weights

Gemma 4 31B

gemma-4-31b-it
Context
256K
Max output
Not published
Input / output
Varies

Capabilities

reasoningfunction callingcodingmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 30.7B dense

License: Apache 2.0

Caveat: Open-weight hosting has no universal price or output cap. The first-party Gemini API deployment can have separate service limits.

Read the AI Wiki article
Official sources (2)
OpenAIActive

GPT-4o

gpt-4o

Aliases: gpt-4o-2024-08-06

Context
128K
Max output
16.4K
Input / output
$2.50 / $10

Capabilities

function callingstructured outputsstreaming

Input: text, image

Output: text

Pricing note: Rates for the gpt-4o alias as listed on October 8, 2026; the pricing page does not say when they took effect. The deprecated gpt-4o-2024-05-13 snapshot bills $5 input and $15 output until it shuts down.

Caveat: Only the older gpt-4o-2024-05-13 snapshot is deprecated, with an October 23, 2026 shutdown. The gpt-4o alias points to gpt-4o-2024-08-06 and is not on OpenAI's deprecation list.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIActive

GPT-5.6 Sol

gpt-5.6-sol

Aliases: gpt-5.6

Context
1.05M
Max output
128K
Input / output
$4 / $20

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Pricing note: Promotional rates are available at least through November 21, 2026. Later rates have not been committed; projections hold these rates constant.

Caveat: The gpt-5.6 alias currently resolves to Sol; aliases are mutable, so pin gpt-5.6-sol when reproducibility matters.

Read the AI Wiki article
Official sources (3)
OpenAIActive

GPT-6 Astra

gpt-6-astra
Context
1.05M
Max output
128K
Input / output
$10 / $50

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Pricing note: Flex is priced like Batch and Fast mode costs 2x Standard. Since September 29, 2026 an Ultrafast tier (service_tier ultrafast, Responses API) costs $60 input, $6 cached input, $75 cache writes and $300 output per 1M tokens, or $120, $12, $150 and $450 above 272K input tokens.

Caveat: Does not accept the none reasoning effort, custom temperature or top_p values, or log probabilities, and tool calling requires the Responses API (Chat Completions works without tools).

Read the AI Wiki article
Official sources (6)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIActive

GPT-6 Luna

gpt-6-luna
Context
1.05M
Max output
128K
Input / output
$0.10 / $0.50

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Pricing note: Flex is priced like Batch, Fast mode costs 2x Standard, and regional processing adds 10 percent where available.

Caveat: On September 25, 2026 OpenAI fixed an image-encoding bug that had degraded image understanding, so rerun image evaluations made before then. Chat Completions supports function calling only with reasoning effort none; use the Responses API for tools with reasoning.

Read the AI Wiki article
Official sources (4)
OpenAIActive

GPT-6 Sol

gpt-6-sol
Context
1.05M
Max output
128K
Input / output
$2 / $10

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Pricing note: Flex is priced like Batch, Fast mode costs 2x Standard, and regional processing adds 10 percent where available.

Caveat: OpenAI's docs now point to GPT-6.1 Sol as the newer Sol model, though gpt-6-sol has no announced deprecation. On September 25, 2026 OpenAI fixed an image-encoding bug that had degraded image understanding, so rerun image evaluations made before then.

Read the AI Wiki article
Official sources (5)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIActive

GPT-6.1 Sol

gpt-6.1-sol
Context
1.05M
Max output
128K
Input / output
$2 / $10

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Pricing note: Cached input is billed at 5 percent of the input rate, half of GPT-6 Sol's cached price. Flex is priced like Batch, Fast mode costs 2x Standard, and regional processing adds 10 percent where available.

Caveat: Does not accept the none or minimal reasoning efforts, and tool calling requires the Responses API (Chat Completions works without tools).

Read the AI Wiki article
Official sources (4)
OpenAIActiveOpen weights

gpt-oss-120b

gpt-oss-120b
Context
131.1K
Max output
131.1K
Input / output
Varies

Capabilities

reasoningfunction callingstructured outputsfine-tuningself-hosting

Input: text

Output: text

Parameters: 117B total / 5.1B active

License: Apache 2.0

Caveat: Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

Read the AI Wiki article
Official sources (3)
OpenAIActiveOpen weights

gpt-oss-20b

gpt-oss-20b
Context
131.1K
Max output
131.1K
Input / output
Varies

Capabilities

reasoningfunction callingstructured outputsfine-tuningself-hosting

Input: text

Output: text

Parameters: 20.9B total / 3.6B active

License: Apache 2.0

Caveat: Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

Read the AI Wiki article
Official sources (3)
xAIActive

Grok 4.3

grok-4.3

Aliases: grok-4.3-latest

Context
1M
Max output
Not published
Input / output
$1.25 / $2.50

Capabilities

reasoningfunction callingstructured outputs

Input: text, image

Output: text

Caveat: The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.

Read the AI Wiki article
Official sources (4)
xAIActive

Grok 4.5

grok-4.5

Aliases: grok-4.5-latest, grok-build-latest

Context
500K
Max output
Not published
Input / output
$2 / $6

Capabilities

reasoningfunction callingstructured outputs

Input: text, image

Output: text

Caveat: The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.

Read the AI Wiki article
Official sources (5)
xAIActive

Grok 4.6

grok-4.6
Context
500K
Max output
Not published
Input / output
$2 / $6

Capabilities

reasoningfunction callingstructured outputsweb searchX searchcode execution

Input: text, image

Output: text

Pricing note: No batch discount: the Batch API is not supported for this model. Requests to the US regional endpoint (https://us.api.x.ai/v1) bill at 1.1x these rates, Priority Processing bills at 2x, and server-side tool calls are billed separately.

Caveat: xAI states that this model has no text output limit and no Batch API support, and xAI publishes no retirement date for it. Its overview page now redirects to Grok 4.7 and the models page recommends Grok 4.7, but grok-4.6 keeps its model page, its pricing row and its place on the US regional endpoint.

Read the AI Wiki article
Official sources (6)
xAIActive

Grok 4.7

grok-4.7
Context
500K
Max output
Not published
Input / output
$2 / $6

Capabilities

reasoningfunction callingstructured outputsweb searchX searchcode execution

Input: text, image

Output: text

Pricing note: No batch discount: the Batch API is not supported for this model. Requests to the US regional endpoint (https://us.api.x.ai/v1) bill at 1.1x these rates, Priority Processing bills at 2x, and server-side tool calls are billed separately.

Caveat: xAI states that this model has no text output limit and no Batch API support, and xAI publishes no retirement date for it. Grok 4.7 Fast (the same model at 2x standard rates, or 1.5x for long-context requests) is sold only in Cursor and Grok Build, not on the public xAI API.

Read the AI Wiki article
Official sources (6)
xAIActive

Grok Build 0.1

grok-build-0.1

Aliases: grok-code-fast-1, grok-code-fast, grok-code-fast-1-0825

Context
256K
Max output
Not published
Input / output
$1 / $2

Capabilities

codingreasoningfunction callingstructured outputs

Input: text, image

Output: text

Caveat: xAI announced grok-build-0.1 on its API in public beta on May 29, 2026 (its May release notes call it early access). The official model page lists no maximum output limit or committed retirement date, and alias mappings can change, so use the explicit model ID.

Read the AI Wiki article
Official sources (6)
MetaActiveOpen weights

Llama 4 Maverick

llama-4-maverick
Context
1M
Max output
Not published
Input / output
Varies

Capabilities

mixture of expertsmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 400B total / 17B active

License: Llama 4 Community License

Caveat: Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers can impose different limits.

Read the AI Wiki article
Official sources (2)
MetaActiveOpen weights

Llama 4 Scout

llama-4-scout
Context
10M
Max output
Not published
Input / output
Varies

Capabilities

mixture of expertsmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 109B total / 17B active

License: Llama 4 Community License

Caveat: Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers may expose a smaller context window.

Read the AI Wiki article
Official sources (2)
MistralActiveAPI + open weights

Mistral Large 3

mistral-large-2512

Aliases: mistral-large-latest

Context
256K
Max output
Not published
Input / output
$0.50 / $1.50

Capabilities

reasoningfunction callingstructured outputsmultilingualself-hosting

Input: text, image

Output: text

Parameters: 675B total / 41B active

License: Apache 2.0

Pricing note: First-party Mistral API price. Cached-input and batch rates are published directly on Mistral's pricing page (Standard and Batch views).

Caveat: mistral-large-latest is a rolling alias that still points to Mistral Large 3 (mistral-large-2512), not to the Mistral Large 4 public preview. The published model card does not specify a separate maximum-output limit.

Read the AI Wiki article
Official sources (7)
MistralActiveAPI + open weights

Mistral Medium 3.5

mistral-medium-3-5

Aliases: mistral-medium-3, mistral-medium-latest

Context
256K
Max output
Not published
Input / output
$1.50 / $7.50

Capabilities

reasoningfunction callingstructured outputscodingself-hosting

Input: text, image

Output: text

License: Modified MIT (per the model card)

Pricing note: First-party Mistral API price. Cached-input and batch rates are published directly on Mistral's pricing page (Standard and Batch views).

Caveat: Mistral publishes dateless IDs only (mistral-medium-3-5 plus the mistral-medium-3 and mistral-medium-latest aliases); the release date and license wording come from the model card alone.

Read the AI Wiki article
Official sources (5)
MistralActiveAPI + open weights

Mistral Small 4

mistral-small-2603

Aliases: mistral-small-latest

Context
256K
Max output
Not published
Input / output
$0.15 / $0.60

Capabilities

reasoningfunction callingstructured outputscodingself-hosting

Input: text, image

Output: text

Parameters: 119B total / 6.5B active

License: Apache 2.0

Pricing note: First-party Mistral API price; self-hosting cost is deployment-specific. Cached-input and batch rates are published directly on Mistral's pricing page (Standard and Batch views).

Caveat: mistral-small-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.

Read the AI Wiki article
Official sources (5)
MetaActive

Muse Spark 1.3

muse-spark-1.3
Context
1.05M
Max output
Not published
Input / output
$1.25 / $4.25

Capabilities

reasoningfunction callingstructured outputsprompt cachingsearch grounding

Input: text, image, video, audio, pdf

Output: text

Pricing note: Standard tier rates as observed on October 8, 2026 (Meta's pricing page does not date them); prompts and completions are not used to train Meta models. No long-context premium; built-in web search grounding adds $2.50 per 1,000 queries. The muse-spark-1.3-contributor variant bills $0.10/$0.002/$0.20 in exchange for training use.

Caveat: Meta lists audio input but says audio understanding in Muse Spark 1.3 is not fully supported and may be degraded (it points audio work to Muse Spark 1.2 or Muse Voice Transcribe). Meta does not publish a maximum output limit, knowledge cutoff, or parameter count, and the max reasoning level is Standard-tier only.

Read the AI Wiki article
Official sources (6)
MetaActive

Muse Spark 1.3 (Contributor tier)

muse-spark-1.3-contributor
Context
1.05M
Max output
Not published
Input / output
$0.10 / $0.20

Capabilities

reasoningfunction callingstructured outputsprompt cachingsearch grounding

Input: text, image, video, audio, pdf

Output: text

Pricing note: Contributor tier rates as observed on October 8, 2026 (Meta's pricing page does not date them): discounted in exchange for permission to use prompts and completions to train future Meta models. The same model on the Standard tier (muse-spark-1.3) costs $1.25/$0.15/$4.25. Built-in web search grounding adds $2.50 per 1,000 queries.

Caveat: Training-eligible tier of Muse Spark 1.3: Meta may use your prompts and completions to train future Meta models, the max reasoning level is not available, and the rate limit is 100 requests and 3,000,000 tokens per minute. Meta does not date this tier's launch, and it says audio understanding in Muse Spark 1.3 is not fully supported.

Read the AI Wiki article
Official sources (5)
OpenAIRetired

GPT-4 32K

gpt-4-32k

Retired

Jun 6, 2025

Past retirement date

Recommended replacement

GPT-4ogpt-4o

OpenAI's documented replacement was gpt-4o, whose alias remains available (only its 2024-05-13 snapshot shuts down on October 23, 2026). OpenAI now recommends the GPT-6 model family for production API use.

Context
32.8K
Max output
Not published
Input / output
Varies

Capabilities

Input: text

Output: text

Caveat: Retired June 6, 2025. Historical pricing is not restated because OpenAI no longer publishes it.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIRetired

OpenAI o1-mini

o1-mini

Retired

Oct 27, 2025

Past retirement date

Announced Apr 28, 2025

Recommended replacement

OpenAI o4-minio4-mini

OpenAI's documented replacement at shutdown was o4-mini, itself deprecated with an October 23, 2026 shutdown. OpenAI now recommends the GPT-6 model family for production API use.

Context
Not published
Max output
Not published
Input / output
Varies

Capabilities

reasoning

Input: text

Output: text

Caveat: Retired October 27, 2025; requests to this model now error. Historical specifications are not restated here.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicRetired

Claude Opus 4

claude-opus-4-20250514

Aliases: claude-opus-4

Retired

Jun 15, 2026

Past retirement date

Announced Apr 14, 2026

Recommended replacement

Claude Opus 4.8claude-opus-4-8

Anthropic's documented replacement is Claude Opus 4.8; requests to the retired ID now return an error on Anthropic-operated platforms.

Context
200K
Max output
32K
Input / output
$15 / $75

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Pricing note: Anthropic still publishes this rate, labeled retired except on Google Cloud.

Caveat: Retired June 15, 2026 on Anthropic-operated platforms; still listed for Google Cloud. Researchers can request access through Anthropic's External Researcher Access Program.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

DeepSeekRetired

deepseek-chat (legacy alias)

deepseek-chat

Retired

Jul 24, 2026

Past retirement date

Announced Apr 24, 2026

Recommended replacement

DeepSeek-V4.1-Flashdeepseek-flash

DeepSeek's original replacement for this alias, deepseek-v4-flash in non-thinking mode, was itself retired on September 10, 2026; that name now routes to deepseek-flash (DeepSeek-V4.1-Flash), which supports both non-thinking and thinking modes.

Context
1M
Max output
384K
Input / output
Varies

Capabilities

Input: text

Output: text

Caveat: Legacy alias with a previously announced July 24, 2026 retirement. Current pricing documentation no longer lists this alias.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

DeepSeekRetired

deepseek-reasoner (legacy alias)

deepseek-reasoner

Retired

Jul 24, 2026

Past retirement date

Announced Apr 24, 2026

Recommended replacement

DeepSeek-V4.1-Flashdeepseek-flash

DeepSeek's original replacement for this alias, deepseek-v4-flash in thinking mode, was itself retired on September 10, 2026; that name now routes to deepseek-flash (DeepSeek-V4.1-Flash), whose default is thinking mode.

Context
1M
Max output
384K
Input / output
Varies

Capabilities

reasoning

Input: text

Output: text

Caveat: Legacy alias with a previously announced July 24, 2026 retirement. Current pricing documentation no longer lists this alias.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicRetired

Claude Opus 4.1

claude-opus-4-1-20250805

Aliases: claude-opus-4-1

Retired

Aug 5, 2026

Past retirement date

Announced Jun 5, 2026

Recommended replacement

Claude Opus 4.8claude-opus-4-8

Audit every dated ID and alias, then test prompting, tool calls, effort settings, output length, and tokenization on Opus 4.8.

Context
200K
Max output
32K
Input / output
$15 / $75

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Caveat: This lifecycle date covers Anthropic-operated platforms. Amazon Bedrock and Google Cloud publish their own schedules.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

DeepSeekRetiredAPI + open weights

DeepSeek-V4-Flash

deepseek-v4-flash

Retired

Sep 10, 2026

Past retirement date

Recommended replacement

DeepSeek-V4.1-Flashdeepseek-flash

The old API name now routes to V4.1-Flash. Use deepseek-flash explicitly.

Context
1M
Max output
384K
Input / output
Varies

Capabilities

reasoningself-hosting

Input: text

Output: text

Parameters: 284B total / 13B active

License: MIT

Caveat: The V4-Flash checkpoint is retired from the first-party API; its downloadable weights remain available. Calls to the old API name are served and billed as V4.1-Flash.

Read the AI Wiki article
Official sources (4)
DeepSeekRetiredAPI + open weights

DeepSeek-V4-Flash-Vision-Exp

deepseek-v4-flash-vision-exp

Retired

Sep 10, 2026

Past retirement date

Announced Sep 10, 2026

Recommended replacement

DeepSeek-V4.1-Flashdeepseek-flash

The old API name now routes to V4.1-Flash and bills at Flash prices. Use deepseek-flash explicitly.

Context
Not published
Max output
Not published
Input / output
Varies

Capabilities

reasoningtool useself-hosting

Input: text, image

Output: text

License: MIT

Caveat: Experimental multimodal build, retired from the first-party API on September 10, 2026; DeepSeek temporarily routes the old name to V4.1-Flash at Flash prices and gives no end date for that routing. The MIT-licensed weights remain on Hugging Face, and DeepSeek's current docs no longer list this model's API context or output limits.

Read the AI Wiki article
Official sources (5)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

Methodology

How to read this tracker

Deprecated means a provider has published a replacement and shutdown milestone. Previewmeans the endpoint is usable but may change on a shorter lifecycle. An availability commitment such as Anthropic's "not sooner than" date is shown separately and is not treated as a retirement announcement.

Prices are standard paid text rates in US dollars per one million tokens. Cached reads and Batch API prices appear in the underlying registry when published, while the cards emphasize the comparable standard input/output pair. Long-context tiers and tool charges are called out but not blended into that number.

Model facts are versioned in the site source so price changes, alias changes, and lifecycle updates can be reviewed rather than silently rewriting history. Every record carries its own verification date and first-party links.

Coverage boundaries

What this does not imply

  • This is a curated set of major general-purpose models, not every image, audio, embedding, or experimental endpoint.
  • API lifecycle dates do not automatically apply to ChatGPT, claude.ai, Gemini apps, or third-party cloud catalogs.
  • "Open-weight" describes downloadable weights; it does not guarantee an open-source license or identical hosted limits.
  • A recommended replacement still needs workload-specific evaluation for quality, safety, latency, tools, and cost.

Source index

Provider pricing and lifecycle notices

These are the core documents used across multiple records. Each model card also exposes its specific model card and release sources.

FAQ

Model lifecycle questions

Are all model retirement dates exact?+

No. OpenAI and Anthropic publish confirmed retirement dates for deprecated API models. Google describes dates in its Gemini deprecation table as the earliest possible shutdown dates and says it will communicate the exact date later. The tracker labels that distinction on every affected record.

Does an API model retirement affect ChatGPT, Claude, or Gemini apps?+

Not necessarily. This tracker follows developer model IDs and first-party API lifecycle notices. Consumer apps can switch their underlying models on a separate schedule, and partner platforms such as Amazon Bedrock or Google Cloud can publish different retirement dates.

What happens when an open-weight model is retired?+

Downloaded weights do not disappear. A hosting provider can retire its endpoint, alias, or managed deployment, but self-hosted checkpoints remain available subject to their license. This is why open-weight records show model facts separately from provider-specific hosting cost.

Can I compare the prices as a complete cost estimate?+

The displayed figures are sourced standard text input and output prices per one million tokens. They are a baseline, not a quote: long-context uplifts, tools, regional processing, priority tiers, cache writes, images, audio, and partner margins can change the total.

Planning an API migration?

Export the lifecycle dates above, inventory every hard-coded model ID and alias, and test the replacement against representative production inputs well before the deadline. For related background, browse large language models, OpenAI, Anthropic, and Google DeepMind.

Machine-readable outputs: JSON feed (CC BY 4.0) and a subscribable calendar (.ics) of confirmed retirement dates.