API migration intelligence

AI model deprecations and retirement tracker

See what is retiring, how long you have, and which model each provider recommends next. Search 56 curated API and open-weight records with release history, context limits, standard text pricing, and links back to official notices.

7 providersOfficial first-party sourcesLatest fact review September 11, 2026; dates vary by fieldCalendar export
Your model workspace

Select models to compare, estimate costs, and check deployment options.

Model lifecycle overview

Next 90 days

9

Retirement or earliest-shutdown milestones in current view.

Deprecated

11

Models in current view that need an active migration plan.

Active

37

Current API and open-weight records in current view.

Open weights

11

Checkpoints in current view; hosting limits and cost vary.

Migration queue

Upcoming lifecycle deadlines

Explore

Find a model or migration

Provider

Showing 56 of 56 curated records

DeepSeekDeprecatedAPI + open weights

DeepSeek-V4-Pro

deepseek-v4-pro

Retirement

Sep 14, 2026, 04:00 UTC

2d 18h remaining (as of September 11)

Recommended replacement

DeepSeek-V4.1-Flashdeepseek-flash

From September 14, 2026 at 04:00 UTC the V4-Pro API name routes to V4.1-Flash and uses Flash pricing.

Context
1M
Max output
384K
Input / output
$1.32 / $3.96

Current pricing; $0.30/$1.20 from Sep 14, 2026.

Capabilities

reasoningself-hosting

Input: text

Output: text

Parameters: 1.6T total / 49B active

License: MIT

Pricing note: Peak rates used for planning. Off-peak input/cache/output: $0.66/$0.022/$1.98. Peak: Monday-Friday 01:00-04:00 and 06:00-10:00 UTC. From September 14 at 04:00 UTC this API name routes to V4.1-Flash; see the Flash record for its rates.

Caveat: 384K is the published maximum output; DeepSeek does not list a default output limit, knowledge cutoff, or deprecation commitments.

Read the AI Wiki article
Official sources (4)
OpenAIDeprecated

GPT-3.5 Turbo

gpt-3.5-turbo

Aliases: gpt-3.5-turbo-0125, gpt-3.5-turbo-completions

Retirement

Oct 23, 2026

41d 14h remaining (as of September 11)

Recommended replacement

GPT-5.6 Terragpt-5.6-terra

Run representative low-latency and high-volume evaluations on Terra and update legacy Chat Completions assumptions.

Context
16.4K
Max output
4.1K
Input / output
$0.50 / $1.50

Capabilities

function callingstreaming

Input: text

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4

gpt-4

Aliases: gpt-4-0613, gpt-4-0613-completions, gpt-4-completions

Retirement

Oct 23, 2026

41d 14h remaining (as of September 11)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-evaluate prompts and tool schemas on GPT-5.6 Sol, then replace every GPT-4 snapshot and Completions alias before shutdown.

Context
8.2K
Max output
8.2K
Input / output
$30 / $60

Capabilities

function callingstreaming

Input: text

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4 Turbo

gpt-4-turbo

Aliases: gpt-4-turbo-2024-04-09, gpt-4-turbo-completions

Retirement

Oct 23, 2026

41d 14h remaining (as of September 11)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Benchmark long prompts and multimodal inputs on GPT-5.6 Sol, then move both the rolling alias and dated snapshot.

Context
128K
Max output
4.1K
Input / output
$10 / $30

Capabilities

function callingstructured outputsstreaming

Input: text, image

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4o

gpt-4o

Aliases: gpt-4o-2024-05-13

Retirement

Oct 23, 2026

41d 14h remaining (as of September 11)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-test multimodal prompts, tool schemas, and latency assumptions on GPT-5.6 Sol before the October 23, 2026 shutdown.

Context
Not published
Max output
Not published
Input / output
Varies

Capabilities

function callingstreaming

Input: text, image

Output: text

Caveat: Lifecycle facts are current; OpenAI no longer publishes pricing or full specifications for this model, so those fields are left unstated rather than guessed.

Read the AI Wiki article
Official sources (1)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o1

o1

Aliases: o1-2024-12-17

Retirement

Oct 23, 2026

41d 14h remaining (as of September 11)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-test reasoning effort, token budgets, and tool behavior on Sol before changing the production model ID.

Context
200K
Max output
100K
Input / output
$15 / $60

Capabilities

reasoningfunction callingstructured outputs

Input: text, image

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o1-pro

o1-pro

Aliases: o1-pro-2025-03-19

Retirement

Oct 23, 2026

41d 14h remaining (as of September 11)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

OpenAI's documented migration is gpt-5.6-sol with reasoning.mode set to pro; re-test reasoning workloads and budgets before cutover.

Context
200K
Max output
100K
Input / output
$150 / $600

Capabilities

reasoningstructured outputs

Input: text, image

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Caveat: Responses API only. OpenAI has not published a launch date; the dated snapshot id suggests March 19, 2025.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o3-mini

o3-mini

Aliases: o3-mini-2025-01-31

Retirement

Oct 23, 2026

41d 14h remaining (as of September 11)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Validate reasoning-intensive workloads on Sol and account for its different price and latency profile.

Context
200K
Max output
100K
Input / output
$1.10 / $4.40

Capabilities

reasoningfunction callingstructured outputs

Input: text

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o4-mini

o4-mini

Aliases: o4-mini-2025-04-16

Retirement

Oct 23, 2026

41d 14h remaining (as of September 11)

Recommended replacement

GPT-5.6 Terragpt-5.6-terra

Test Terra with the same tools and reasoning workloads; check quality, latency, and token use before cutover.

Context
200K
Max output
100K
Input / output
$1.10 / $4.40

Capabilities

reasoningfunction callingstructured outputs

Input: text, image

Output: text

Pricing note: Published on the individual model reference; this deprecated model is omitted from the main pricing table.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleDeprecated

Gemini 3.1 Flash-Lite

gemini-3.1-flash-lite

Earliest shutdown

May 7, 2027

7 months remaining (as of September 11)

Recommended replacement

Gemini 3.5 Flash-Litegemini-3.5-flash-lite

Google names Gemini 3.5 Flash-Lite as the successor; compare quality and latency before moving latency-sensitive workloads.

Context
1.05M
Max output
65.5K
Input / output
$0.25 / $1.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; audio input is billed at $0.50 per 1M. Context caching adds $1.00 per 1M tokens per hour of storage.

Caveat: Google lists May 7, 2027 as the earliest possible shutdown while this model remains the documented migration target for Gemini 2.5 Flash-Lite.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicDeprecated

Claude Mythos Preview

claude-mythos-preview

Recommended replacement

Claude Mythos 5claude-mythos-5

Anthropic's documented migration target is Claude Mythos 5, which requires Project Glasswing approval.

Context
1M
Max output
Not published
Input / output
Varies

Capabilities

Input: text, image

Output: text

Caveat: The current lifecycle documentation marks Mythos Preview deprecated and directs users to Mythos 5. It does not provide a retirement date; do not treat the former July date as a current commitment.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GooglePreview

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview

Aliases: gemini-pro-latest

Context
1.05M
Max output
65.5K
Input / output
$2 / $12

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Caveat: Preview models can change or retire on shorter notice. gemini-pro-latest is a rolling alias, not a pinned release.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Fable 5

claude-fable-5

Availability commitment

No retirement sooner than Jun 9, 2027

Context
1M
Max output
128K
Input / output
$10 / $50

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Caveat: US export controls forced a global suspension of Fable 5 and Mythos 5 from June 12 to June 30, 2026; access was restored on July 1, 2026. Availability and data-retention eligibility can vary.

Read the AI Wiki article
Official sources (5)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Haiku 4.5

claude-haiku-4-5-20251001

Aliases: claude-haiku-4-5

Availability commitment

No retirement sooner than Oct 15, 2026

Context
200K
Max output
64K
Input / output
$1 / $5

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Caveat: The dated API ID is pinned. The shorter claude-haiku-4-5 value is a convenience alias.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Mythos 5

claude-mythos-5
Context
1M
Max output
128K
Input / output
$10 / $50

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Anthropic lists the same specs and pricing as Claude Fable 5.

Caveat: Not generally available: offered in limited availability to approved Project Glasswing customers for defensive cybersecurity work. Also covered by the June 12 to June 30, 2026 export-control suspension; access was restored July 1, 2026.

Read the AI Wiki article
Official sources (3)
AnthropicActive

Claude Opus 4.5

claude-opus-4-5-20251101

Aliases: claude-opus-4-5

Availability commitment

No retirement sooner than Nov 24, 2026

Context
200K
Max output
64K
Input / output
$5 / $25

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Opus 4.6

claude-opus-4-6

Availability commitment

No retirement sooner than Feb 5, 2027

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Opus 4.7

claude-opus-4-7

Availability commitment

No retirement sooner than Apr 16, 2027

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models. Fast mode was deprecated June 25, 2026, with removal on July 24, 2026.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Opus 4.8

claude-opus-4-8

Availability commitment

No retirement sooner than May 28, 2027

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Sonnet 4.5

claude-sonnet-4-5-20250929

Aliases: claude-sonnet-4-5

Availability commitment

No retirement sooner than Sep 29, 2026

Context
200K
Max output
64K
Input / output
$3 / $15

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Sonnet 4.6

claude-sonnet-4-6

Availability commitment

No retirement sooner than Feb 17, 2027

Context
1M
Max output
128K
Input / output
$3 / $15

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Sonnet 5

claude-sonnet-5

Availability commitment

No retirement sooner than Jun 30, 2027

Context
1M
Max output
128K
Input / output
$2 / $10

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: These are standard prices. The previously announced September 1 increase was cancelled.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

DeepSeekActiveAPI + open weights

DeepSeek-V4.1-Flash

deepseek-flash
Context
1M
Max output
384K
Input / output
$0.30 / $1.20

Capabilities

reasoningtool usestructured outputsself-hosting

Input: text, image

Output: text

Parameters: 552B backbone + 196B Engram; 8B active prefill / 16B decode

License: MIT

Pricing note: Peak rates used for planning. Off-peak input/cache/output: $0.15/$0.003/$0.60. Peak hours: Monday-Friday 01:00-04:00 and 06:00-10:00 UTC; all other hours are half price.

Caveat: API model ID is deepseek-flash. Generic dense-transformer memory formulas do not model this architecture accurately.

Read the AI Wiki article
Official sources (2)
GoogleActive

Gemini 2.5 Flash

gemini-2.5-flash
Context
1.05M
Max output
65.5K
Input / output
$0.30 / $2.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Caveat: The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 2.5 Flash-Lite

gemini-2.5-flash-lite
Context
1.05M
Max output
65.5K
Input / output
$0.10 / $0.40

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Caveat: The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 2.5 Pro

gemini-2.5-pro
Context
1.05M
Max output
65.5K
Input / output
$1.25 / $10

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Caveat: The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.5 Flash

gemini-3.5-flash

Aliases: gemini-flash-latest

Context
1.05M
Max output
65.5K
Input / output
$1.50 / $9

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; feature charges are separate.

Caveat: The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.5 Flash-Lite

gemini-3.5-flash-lite
Context
1.05M
Max output
65.5K
Input / output
$0.30 / $2.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; feature charges are separate. Context caching is paid-tier only (plus $1.00 per 1M tokens per hour of cache storage).

Caveat: The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.6 Flash

gemini-3.6-flash
Context
1.05M
Max output
65.5K
Input / output
$0.75 / $3.75

Current pricing through Dec 31, 2026; $1.50/$7.50 from Jan 1, 2027.

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Rates observed September 11; the pricing page does not state when the discount began. Cache storage and grounding are additional charges.

Caveat: The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.8 Flash

gemini-3.8-flash
Context
1.05M
Max output
65.5K
Input / output
$0.75 / $3.75

Current pricing through Dec 31, 2026; $1.50/$7.50 from Jan 1, 2027.

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingbatch processing

Input: text, image, audio, video, pdf

Output: text

Pricing note: Rates observed September 11; the pricing page does not state when the discount began. Cache storage and grounding are additional charges.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActiveAPI + open weights

Gemma 4 31B

gemma-4-31b-it
Context
256K
Max output
Not published
Input / output
Varies

Capabilities

reasoningfunction callingcodingmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 30.7B dense

License: Apache 2.0

Caveat: Open-weight hosting has no universal price or output cap. The first-party Gemini API deployment can have separate service limits.

Read the AI Wiki article
Official sources (2)
OpenAIActive

GPT-5.6 Sol

gpt-5.6-sol

Aliases: gpt-5.6

Context
1.05M
Max output
128K
Input / output
$4 / $20

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Pricing note: Promotional rates are available at least through November 21, 2026. Later rates have not been committed; projections hold these rates constant.

Caveat: The gpt-5.6 alias currently resolves to Sol; aliases are mutable, so pin gpt-5.6-sol when reproducibility matters.

Read the AI Wiki article
Official sources (3)
OpenAIActiveOpen weights

gpt-oss-120b

gpt-oss-120b
Context
131.1K
Max output
131.1K
Input / output
Varies

Capabilities

reasoningfunction callingstructured outputsfine-tuningself-hosting

Input: text

Output: text

Parameters: 117B total / 5.1B active

License: Apache 2.0

Caveat: Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

Read the AI Wiki article
Official sources (3)
OpenAIActiveOpen weights

gpt-oss-20b

gpt-oss-20b
Context
131.1K
Max output
131.1K
Input / output
Varies

Capabilities

reasoningfunction callingstructured outputsfine-tuningself-hosting

Input: text

Output: text

Parameters: 20.9B total / 3.6B active

License: Apache 2.0

Caveat: Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

Read the AI Wiki article
Official sources (3)
xAIActive

Grok 4.3

grok-4.3

Aliases: grok-4.3-latest

Context
1M
Max output
Not published
Input / output
$1.25 / $2.50

Capabilities

Input: text, image

Output: text

Caveat: The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.

Read the AI Wiki article
Official sources (4)
xAIActive

Grok 4.5

grok-4.5

Aliases: grok-4.5-latest, grok-build-latest

Context
500K
Max output
Not published
Input / output
$2 / $6

Capabilities

Input: text, image

Output: text

Caveat: The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.

Read the AI Wiki article
Official sources (5)
xAIActive

Grok Build 0.1

grok-build-0.1

Aliases: grok-code-fast-1, grok-code-fast, grok-code-fast-1-0825

Context
256K
Max output
Not published
Input / output
$1 / $2

Capabilities

coding

Input: text, image

Output: text

Caveat: The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.

Read the AI Wiki article
Official sources (4)
MetaActiveOpen weights

Llama 4 Maverick

llama-4-maverick
Context
1M
Max output
Not published
Input / output
Varies

Capabilities

mixture of expertsmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 400B total / 17B active

License: Llama 4 Community License

Caveat: Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers can impose different limits.

Read the AI Wiki article
Official sources (2)
MetaActiveOpen weights

Llama 4 Scout

llama-4-scout
Context
10M
Max output
Not published
Input / output
Varies

Capabilities

mixture of expertsmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 109B total / 17B active

License: Llama 4 Community License

Caveat: Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers may expose a smaller context window.

Read the AI Wiki article
Official sources (2)
MistralActiveAPI + open weights

Mistral Large 3

mistral-large-2512

Aliases: mistral-large-latest

Context
256K
Max output
Not published
Input / output
$0.50 / $1.50

Capabilities

reasoningfunction callingstructured outputsmultilingualself-hosting

Input: text, image

Output: text

Parameters: 675B total / 41B active

License: Apache 2.0

Pricing note: First-party Mistral API price. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.

Caveat: mistral-large-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.

Read the AI Wiki article
Official sources (5)
MistralActiveAPI + open weights

Mistral Medium 3.5

mistral-medium-3-5

Aliases: mistral-medium-3, mistral-medium-latest

Context
256K
Max output
Not published
Input / output
$1.50 / $7.50

Capabilities

reasoningfunction callingstructured outputscodingself-hosting

Input: text, image

Output: text

License: Modified MIT (per the model card)

Pricing note: First-party Mistral API price. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.

Caveat: Mistral publishes dateless IDs only (mistral-medium-3-5 plus the mistral-medium-3 and mistral-medium-latest aliases); the release date and license wording come from the model card alone.

Read the AI Wiki article
Official sources (4)
MistralActiveAPI + open weights

Mistral Small 4

mistral-small-2603

Aliases: mistral-small-latest

Context
256K
Max output
Not published
Input / output
$0.15 / $0.60

Capabilities

reasoningfunction callingstructured outputscodingself-hosting

Input: text, image

Output: text

Parameters: 119B total / 6.5B active

License: Apache 2.0

Pricing note: First-party Mistral API price; self-hosting cost is deployment-specific. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.

Caveat: mistral-small-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.

Read the AI Wiki article
Official sources (4)
OpenAIRetired

GPT-4 32K

gpt-4-32k

Retired

Jun 6, 2025

Past retirement date

Recommended replacement

GPT-4ogpt-4o

OpenAI pointed gpt-4-32k users to gpt-4o, which is itself deprecated with an October 23, 2026 shutdown; new work belongs on the GPT-5.6 family.

Context
32.8K
Max output
Not published
Input / output
Varies

Capabilities

Input: text

Output: text

Caveat: Retired June 6, 2025. Historical pricing is not restated because OpenAI no longer publishes it.

Read the AI Wiki article
Official sources (1)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIRetired

OpenAI o1-mini

o1-mini

Retired

Oct 27, 2025

Past retirement date

Announced Apr 28, 2025

Recommended replacement

OpenAI o4-minio4-mini

OpenAI's documented replacement at shutdown was o4-mini, itself now deprecated with an October 23, 2026 shutdown; new work belongs on the GPT-5.6 family.

Context
Not published
Max output
Not published
Input / output
Varies

Capabilities

reasoning

Input: text

Output: text

Caveat: Retired October 27, 2025; requests to this model now error. Historical specifications are not restated here.

Read the AI Wiki article
Official sources (1)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicRetired

Claude Opus 4

claude-opus-4-20250514

Aliases: claude-opus-4

Retired

Jun 15, 2026

Past retirement date

Announced Apr 14, 2026

Recommended replacement

Claude Opus 4.8claude-opus-4-8

Anthropic's documented replacement is Claude Opus 4.8; requests to the retired ID now return an error on Anthropic-operated platforms.

Context
200K
Max output
32K
Input / output
$15 / $75

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Pricing note: Anthropic still publishes this rate, labeled retired except on Google Cloud.

Caveat: Retired June 15, 2026 on Anthropic-operated platforms; still listed for Google Cloud. Researchers can request access through Anthropic's External Researcher Access Program.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

DeepSeekRetired

deepseek-chat (legacy alias)

deepseek-chat

Retired

Jul 24, 2026

Past retirement date

Announced Apr 24, 2026

Recommended replacement

DeepSeek-V4-Flashdeepseek-v4-flash

Point API calls at deepseek-v4-flash; its non-thinking mode matches the old deepseek-chat behavior.

Context
1M
Max output
384K
Input / output
Varies

Capabilities

Input: text

Output: text

Caveat: Legacy alias with a previously announced July 24, 2026 retirement. Current pricing documentation no longer lists this alias.

Read the AI Wiki article
Official sources (2)
DeepSeekRetired

deepseek-reasoner (legacy alias)

deepseek-reasoner

Retired

Jul 24, 2026

Past retirement date

Announced Apr 24, 2026

Recommended replacement

DeepSeek-V4-Flashdeepseek-v4-flash

Point API calls at deepseek-v4-flash; its thinking mode matches the old deepseek-reasoner behavior.

Context
1M
Max output
384K
Input / output
Varies

Capabilities

reasoning

Input: text

Output: text

Caveat: Legacy alias with a previously announced July 24, 2026 retirement. Current pricing documentation no longer lists this alias.

Read the AI Wiki article
Official sources (2)
AnthropicRetired

Claude Opus 4.1

claude-opus-4-1-20250805

Aliases: claude-opus-4-1

Retired

Aug 5, 2026

Past retirement date

Announced Jun 5, 2026

Recommended replacement

Claude Opus 4.8claude-opus-4-8

Audit every dated ID and alias, then test prompting, tool calls, effort settings, output length, and tokenization on Opus 4.8.

Context
200K
Max output
32K
Input / output
$15 / $75

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Caveat: This lifecycle date covers Anthropic-operated platforms. Amazon Bedrock and Google Cloud publish their own schedules.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

DeepSeekRetiredAPI + open weights

DeepSeek-V4-Flash

deepseek-v4-flash

Retired

Sep 10, 2026

Past retirement date

Recommended replacement

DeepSeek-V4.1-Flashdeepseek-flash

The old API name now routes to V4.1-Flash. Use deepseek-flash explicitly.

Context
1M
Max output
384K
Input / output
Varies

Capabilities

reasoningself-hosting

Input: text

Output: text

Parameters: 284B total / 13B active

License: MIT

Caveat: The V4-Flash checkpoint is retired from the first-party API; its downloadable weights remain available. Calls to the old API name are served and billed as V4.1-Flash.

Read the AI Wiki article
Official sources (4)

Methodology

How to read this tracker

Deprecated means a provider has published a replacement and shutdown milestone. Previewmeans the endpoint is usable but may change on a shorter lifecycle. An availability commitment such as Anthropic's "not sooner than" date is shown separately and is not treated as a retirement announcement.

Prices are standard paid text rates in US dollars per one million tokens. Cached reads and Batch API prices appear in the underlying registry when published, while the cards emphasize the comparable standard input/output pair. Long-context tiers and tool charges are called out but not blended into that number.

Model facts are versioned in the site source so price changes, alias changes, and lifecycle updates can be reviewed rather than silently rewriting history. Every record carries its own verification date and first-party links.

Coverage boundaries

What this does not imply

  • This is a curated set of major general-purpose models, not every image, audio, embedding, or experimental endpoint.
  • API lifecycle dates do not automatically apply to ChatGPT, claude.ai, Gemini apps, or third-party cloud catalogs.
  • "Open-weight" describes downloadable weights; it does not guarantee an open-source license or identical hosted limits.
  • A recommended replacement still needs workload-specific evaluation for quality, safety, latency, tools, and cost.

Source index

Provider pricing and lifecycle notices

These are the core documents used across multiple records. Each model card also exposes its specific model card and release sources.

FAQ

Model lifecycle questions

Are all model retirement dates exact?+

No. OpenAI and Anthropic publish confirmed retirement dates for deprecated API models. Google describes dates in its Gemini deprecation table as the earliest possible shutdown dates and says it will communicate the exact date later. The tracker labels that distinction on every affected record.

Does an API model retirement affect ChatGPT, Claude, or Gemini apps?+

Not necessarily. This tracker follows developer model IDs and first-party API lifecycle notices. Consumer apps can switch their underlying models on a separate schedule, and partner platforms such as Amazon Bedrock or Google Cloud can publish different retirement dates.

What happens when an open-weight model is retired?+

Downloaded weights do not disappear. A hosting provider can retire its endpoint, alias, or managed deployment, but self-hosted checkpoints remain available subject to their license. This is why open-weight records show model facts separately from provider-specific hosting cost.

Can I compare the prices as a complete cost estimate?+

The displayed figures are sourced standard text input and output prices per one million tokens. They are a baseline, not a quote: long-context uplifts, tools, regional processing, priority tiers, cache writes, images, audio, and partner margins can change the total.

Planning an API migration?

Export the lifecycle dates above, inventory every hard-coded model ID and alias, and test the replacement against representative production inputs well before the deadline. For related background, browse large language models, OpenAI, Anthropic, and Google DeepMind.

Machine-readable outputs: JSON feed (CC BY 4.0) and a subscribable calendar (.ics) of confirmed retirement dates.