API migration intelligence

AI model deprecations, without the deadline panic

See what is retiring, how long you have, and which model each provider recommends next. Search 53 curated API and open-weight records with release history, context limits, standard text pricing, and links back to official notices.

7 providersOfficial first-party sourcesReviewed July 23, 2026Calendar export

Model lifecycle overview

Next 90 days

6

Retirement or earliest-shutdown milestones in current view.

Deprecated

15

Models in current view that need an active migration plan.

Active

28

Current API and open-weight records in current view.

Open weights

10

Checkpoints in current view; hosting limits and cost vary.

Migration queue

Upcoming lifecycle deadlines

Explore

Find a model or migration

Provider

Showing 53 of 53 curated records

DeepSeekDeprecated

deepseek-chat (legacy alias)

deepseek-chat

Retirement

Jul 24, 2026

Past retirement date

Announced Apr 24, 2026

Recommended replacement

DeepSeek-V4-Flashdeepseek-v4-flash

Point API calls at deepseek-v4-flash; its non-thinking mode matches the old deepseek-chat behavior.

Context
1M
Max output
384K
Input / output
Varies

Capabilities

Input: text

Output: text

Caveat: Model-name alias retiring at 15:59 UTC on July 24, 2026; it currently resolves to deepseek-v4-flash in non-thinking mode and bills at that model's rates.

Read the AI Wiki article
Official sources (2)
DeepSeekDeprecated

deepseek-reasoner (legacy alias)

deepseek-reasoner

Retirement

Jul 24, 2026

Past retirement date

Announced Apr 24, 2026

Recommended replacement

DeepSeek-V4-Flashdeepseek-v4-flash

Point API calls at deepseek-v4-flash; its thinking mode matches the old deepseek-reasoner behavior.

Context
1M
Max output
384K
Input / output
Varies

Capabilities

reasoning

Input: text

Output: text

Caveat: Model-name alias retiring at 15:59 UTC on July 24, 2026; it currently resolves to deepseek-v4-flash in thinking mode and bills at that model's rates.

Read the AI Wiki article
Official sources (2)
AnthropicDeprecated

Claude Opus 4.1

claude-opus-4-1-20250805

Aliases: claude-opus-4-1

Retirement

Aug 5, 2026

11d 15h remaining (as of July 24)

Announced Jun 5, 2026

Recommended replacement

Claude Opus 4.8claude-opus-4-8

Audit every dated ID and alias, then test prompting, tool calls, effort settings, output length, and tokenization on Opus 4.8.

Context
200K
Max output
32K
Input / output
$15 / $75

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Caveat: This lifecycle date covers Anthropic-operated platforms. Amazon Bedrock and Google Cloud publish their own schedules.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleDeprecated

Gemini 2.5 Flash

gemini-2.5-flash

Earliest shutdown

Oct 16, 2026

83d 15h remaining (as of July 24)

Recommended replacement

Gemini 3.6 Flashgemini-3.6-flash

Compare thinking behavior, latency, tool calls, and multimodal prompts on the stable 3.6 endpoint before switching.

Context
1.05M
Max output
65.5K
Input / output
$0.30 / $2.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Caveat: Google labels October 16 as the earliest possible shutdown, not a guaranteed exact shutdown day.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleDeprecated

Gemini 2.5 Flash-Lite

gemini-2.5-flash-lite

Earliest shutdown

Oct 16, 2026

83d 15h remaining (as of July 24)

Recommended replacement

Gemini 3.1 Flash-Litegemini-3.1-flash-lite

Google's documented target is Gemini 3.1 Flash-Lite, which itself carries an earliest shutdown of May 7, 2027 with Gemini 3.5 Flash-Lite named as its successor; evaluate 3.5 Flash-Lite directly for new work.

Context
1.05M
Max output
65.5K
Input / output
$0.10 / $0.40

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Caveat: Google labels October 16 as the earliest possible shutdown, not a guaranteed exact shutdown day.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleDeprecated

Gemini 2.5 Pro

gemini-2.5-pro

Earliest shutdown

Oct 16, 2026

83d 15h remaining (as of July 24)

Recommended replacement

Gemini 3.1 Pro Previewgemini-3.1-pro-preview

Google currently recommends a preview replacement. Test thought signatures, tools, structured output, and long prompts before migrating.

Context
1.05M
Max output
65.5K
Input / output
$1.25 / $10

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Caveat: Google labels October 16 as the earliest possible shutdown, not a guaranteed exact shutdown day.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-3.5 Turbo

gpt-3.5-turbo

Aliases: gpt-3.5-turbo-0125, gpt-3.5-turbo-completions

Retirement

Oct 23, 2026

3 months remaining (as of July 24)

Recommended replacement

GPT-5.6 Terragpt-5.6-terra

Run representative low-latency and high-volume evaluations on Terra and update legacy Chat Completions assumptions.

Context
16.4K
Max output
4.1K
Input / output
$0.50 / $1.50

Capabilities

function callingstreaming

Input: text

Output: text

Pricing note: Last published list price. OpenAI's current pricing page no longer lists this deprecated model.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4

gpt-4

Aliases: gpt-4-0613, gpt-4-0613-completions, gpt-4-completions

Retirement

Oct 23, 2026

3 months remaining (as of July 24)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-evaluate prompts and tool schemas on GPT-5.6 Sol, then replace every GPT-4 snapshot and Completions alias before shutdown.

Context
8.2K
Max output
8.2K
Input / output
$30 / $60

Capabilities

function callingstreaming

Input: text

Output: text

Pricing note: Last published list price. OpenAI's current pricing page no longer lists this deprecated model.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4 Turbo

gpt-4-turbo

Aliases: gpt-4-turbo-2024-04-09, gpt-4-turbo-completions

Retirement

Oct 23, 2026

3 months remaining (as of July 24)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Benchmark long prompts and multimodal inputs on GPT-5.6 Sol, then move both the rolling alias and dated snapshot.

Context
128K
Max output
4.1K
Input / output
$10 / $30

Capabilities

function callingstructured outputsstreaming

Input: text, image

Output: text

Pricing note: Last published list price. OpenAI's current pricing page no longer lists this deprecated model.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

GPT-4o

gpt-4o

Aliases: gpt-4o-2024-05-13

Retirement

Oct 23, 2026

3 months remaining (as of July 24)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-test multimodal prompts, tool schemas, and latency assumptions on GPT-5.6 Sol before the October 23, 2026 shutdown.

Context
Not published
Max output
Not published
Input / output
Varies

Capabilities

function callingstreaming

Input: text, image

Output: text

Caveat: Lifecycle facts are current; OpenAI no longer publishes pricing or full specifications for this model, so those fields are left unstated rather than guessed.

Read the AI Wiki article
Official sources (1)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o1

o1

Aliases: o1-2024-12-17

Retirement

Oct 23, 2026

3 months remaining (as of July 24)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Re-test reasoning effort, token budgets, and tool behavior on Sol before changing the production model ID.

Context
200K
Max output
100K
Input / output
$15 / $60

Capabilities

reasoningfunction callingstructured outputs

Input: text, image

Output: text

Pricing note: Last published list price. OpenAI's current pricing page no longer lists this deprecated model.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o1-pro

o1-pro

Aliases: o1-pro-2025-03-19

Retirement

Oct 23, 2026

3 months remaining (as of July 24)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

OpenAI's documented migration is gpt-5.6-sol with reasoning.mode set to pro; re-test reasoning workloads and budgets before cutover.

Context
200K
Max output
100K
Input / output
$150 / $600

Capabilities

reasoningstructured outputs

Input: text, image

Output: text

Pricing note: Published on the o1-pro model page; absent from OpenAI's main pricing page.

Caveat: Responses API only. OpenAI has not published a launch date; the dated snapshot id suggests March 19, 2025.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o3-mini

o3-mini

Aliases: o3-mini-2025-01-31

Retirement

Oct 23, 2026

3 months remaining (as of July 24)

Recommended replacement

GPT-5.6 Solgpt-5.6-sol

Validate reasoning-intensive workloads on Sol and account for its different price and latency profile.

Context
200K
Max output
100K
Input / output
$1.10 / $4.40

Capabilities

reasoningfunction callingstructured outputs

Input: text

Output: text

Pricing note: Last published list price. OpenAI's current pricing page no longer lists this deprecated model.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIDeprecated

OpenAI o4-mini

o4-mini

Aliases: o4-mini-2025-04-16

Retirement

Oct 23, 2026

3 months remaining (as of July 24)

Recommended replacement

GPT-5.6 Terragpt-5.6-terra

Test Terra with the same tools and reasoning workloads; check quality, latency, and token use before cutover.

Context
200K
Max output
100K
Input / output
$1.10 / $4.40

Capabilities

reasoningfunction callingstructured outputs

Input: text, image

Output: text

Pricing note: Last published list price. OpenAI's current pricing page no longer lists this deprecated model.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleDeprecated

Gemini 3.1 Flash-Lite

gemini-3.1-flash-lite

Earliest shutdown

May 7, 2027

9 months remaining (as of July 24)

Recommended replacement

Gemini 3.5 Flash-Litegemini-3.5-flash-lite

Google names Gemini 3.5 Flash-Lite as the successor; compare quality and latency before moving latency-sensitive workloads.

Context
1.05M
Max output
65.5K
Input / output
$0.25 / $1.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; audio input is billed at $0.50 per 1M. Context caching adds $1.00 per 1M tokens per hour of storage.

Caveat: Google lists May 7, 2027 as the earliest possible shutdown while this model remains the documented migration target for Gemini 2.5 Flash-Lite.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicLegacy

Claude Opus 4.5

claude-opus-4-5-20251101

Aliases: claude-opus-4-5

Legacy: availability commitment

No retirement sooner than Nov 24, 2026

Grouped under legacy models but not deprecated; no shutdown date has been announced.

Context
200K
Max output
64K
Input / output
$5 / $25

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicLegacy

Claude Opus 4.6

claude-opus-4-6

Legacy: availability commitment

No retirement sooner than Feb 5, 2027

Grouped under legacy models but not deprecated; no shutdown date has been announced.

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicLegacy

Claude Opus 4.7

claude-opus-4-7

Legacy: availability commitment

No retirement sooner than Apr 16, 2027

Grouped under legacy models but not deprecated; no shutdown date has been announced.

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models. Fast mode was deprecated June 25, 2026, with removal on July 24, 2026.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicLegacy

Claude Sonnet 4.5

claude-sonnet-4-5-20250929

Aliases: claude-sonnet-4-5

Legacy: availability commitment

No retirement sooner than Sep 29, 2026

Grouped under legacy models but not deprecated; no shutdown date has been announced.

Context
200K
Max output
64K
Input / output
$3 / $15

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicLegacy

Claude Sonnet 4.6

claude-sonnet-4-6

Legacy: availability commitment

No retirement sooner than Feb 17, 2027

Grouped under legacy models but not deprecated; no shutdown date has been announced.

Context
1M
Max output
128K
Input / output
$3 / $15

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Caveat: Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GooglePreview

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview

Aliases: gemini-pro-latest

Context
1.05M
Max output
65.5K
Input / output
$2 / $12

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.

Caveat: Preview models can change or retire on shorter notice. gemini-pro-latest is a rolling alias, not a pinned release.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Fable 5

claude-fable-5

Availability commitment

No retirement sooner than Jun 9, 2027

Context
1M
Max output
128K
Input / output
$10 / $50

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Caveat: US export controls forced a global suspension of Fable 5 and Mythos 5 from June 12 to June 30, 2026; access was restored on July 1, 2026. Availability and data-retention eligibility can vary.

Read the AI Wiki article
Official sources (5)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Haiku 4.5

claude-haiku-4-5-20251001

Aliases: claude-haiku-4-5

Availability commitment

No retirement sooner than Oct 15, 2026

Context
200K
Max output
64K
Input / output
$1 / $5

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Caveat: The dated API ID is pinned. The shorter claude-haiku-4-5 value is a convenience alias.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Mythos 5

claude-mythos-5
Context
1M
Max output
128K
Input / output
$10 / $50

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Anthropic lists the same specs and pricing as Claude Fable 5.

Caveat: Not generally available: offered in limited availability to approved Project Glasswing customers for defensive cybersecurity work. Also covered by the June 12 to June 30, 2026 export-control suspension; access was restored July 1, 2026.

Read the AI Wiki article
Official sources (3)
AnthropicActive

Claude Opus 4.8

claude-opus-4-8

Availability commitment

No retirement sooner than May 28, 2027

Context
1M
Max output
128K
Input / output
$5 / $25

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Base Claude API price; regional and partner-platform charges can differ.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicActive

Claude Sonnet 5

claude-sonnet-5

Availability commitment

No retirement sooner than Jun 30, 2027

Context
1M
Max output
128K
Input / output
$2 / $10

Intro pricing through Aug 31, 2026; $3/$15 from Sep 1, 2026.

Capabilities

adaptive thinkingvisiontool useprompt cachingbatch processing

Input: text, image, pdf

Output: text

Pricing note: Introductory pricing through August 31, 2026; Anthropic lists $3 input / $15 output afterward.

Read the AI Wiki article
Official sources (4)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

DeepSeekActiveAPI + open weights

DeepSeek-V4-Flash

deepseek-v4-flash
Context
1M
Max output
384K
Input / output
$0.14 / $0.28

Capabilities

reasoningself-hosting

Input: text

Output: text

Parameters: 284B total / 13B active

License: MIT

Pricing note: Cache-miss input price; cache hits bill at $0.0028 per 1M. One price covers thinking and non-thinking modes.

Caveat: 384K is the published maximum output; DeepSeek does not list a default output limit, knowledge cutoff, or deprecation commitments.

Read the AI Wiki article
Official sources (3)
DeepSeekActiveAPI + open weights

DeepSeek-V4-Pro

deepseek-v4-pro
Context
1M
Max output
384K
Input / output
$0.43 / $0.87

Capabilities

reasoningself-hosting

Input: text

Output: text

Parameters: 1.6T total / 49B active

License: MIT

Pricing note: Cache-miss input price; cache hits bill at $0.003625 per 1M. One price covers thinking and non-thinking modes.

Caveat: 384K is the published maximum output; DeepSeek does not list a default output limit, knowledge cutoff, or deprecation commitments.

Read the AI Wiki article
Official sources (3)
GoogleActive

Gemini 3.5 Flash

gemini-3.5-flash

Aliases: gemini-flash-latest

Context
1.05M
Max output
65.5K
Input / output
$1.50 / $9

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; feature charges are separate.

Caveat: gemini-flash-latest is a rolling alias and can change target. Use the stable model ID for production pinning.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.5 Flash-Lite

gemini-3.5-flash-lite
Context
1.05M
Max output
65.5K
Input / output
$0.30 / $2.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; feature charges are separate. Context caching is paid-tier only (plus $1.00 per 1M tokens per hour of cache storage).

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActive

Gemini 3.6 Flash

gemini-3.6-flash
Context
1.05M
Max output
65.5K
Input / output
$1.50 / $7.50

Capabilities

thinkingfunction callingstructured outputscode executionsearch groundingcontext caching

Input: text, image, video, audio, pdf

Output: text

Pricing note: Paid-tier text/image/video input price; feature charges are separate.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

GoogleActiveAPI + open weights

Gemma 4 31B

gemma-4-31b-it
Context
256K
Max output
Not published
Input / output
Varies

Capabilities

reasoningfunction callingcodingmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 30.7B dense

License: Apache 2.0

Caveat: Open-weight hosting has no universal price or output cap. The first-party Gemini API deployment can have separate service limits.

Read the AI Wiki article
Official sources (2)
OpenAIActive

GPT-5.6 Luna

gpt-5.6-luna
Context
1.05M
Max output
128K
Input / output
$1 / $6

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Pricing note: Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.

Read the AI Wiki article
Official sources (3)
OpenAIActive

GPT-5.6 Sol

gpt-5.6-sol

Aliases: gpt-5.6

Context
1.05M
Max output
128K
Input / output
$5 / $30

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Pricing note: Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.

Caveat: The gpt-5.6 alias currently resolves to Sol; aliases are mutable, so pin gpt-5.6-sol when reproducibility matters.

Read the AI Wiki article
Official sources (3)
OpenAIActive

GPT-5.6 Terra

gpt-5.6-terra
Context
1.05M
Max output
128K
Input / output
$2.50 / $15

Capabilities

reasoningfunction callingstructured outputsstreamingtool use

Input: text, image

Output: text

Pricing note: Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.

Read the AI Wiki article
Official sources (3)
OpenAIActiveOpen weights

gpt-oss-120b

gpt-oss-120b
Context
131.1K
Max output
131.1K
Input / output
Varies

Capabilities

reasoningfunction callingstructured outputsfine-tuningself-hosting

Input: text

Output: text

Parameters: 117B total / 5.1B active

License: Apache 2.0

Caveat: Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

Read the AI Wiki article
Official sources (3)
OpenAIActiveOpen weights

gpt-oss-20b

gpt-oss-20b
Context
131.1K
Max output
131.1K
Input / output
Varies

Capabilities

reasoningfunction callingstructured outputsfine-tuningself-hosting

Input: text

Output: text

Parameters: 20.9B total / 3.6B active

License: Apache 2.0

Caveat: Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.

Read the AI Wiki article
Official sources (3)
xAIActive

Grok Build 0.1

grok-build-0.1

Aliases: grok-code-fast-1, grok-code-fast, grok-code-fast-1-0825

Context
256K
Max output
Not published
Input / output
$1 / $2

Capabilities

coding

Input: text, image

Output: text

Caveat: Coding-focused model; legacy grok-code-fast IDs resolve here per xAI's docs. xAI does not publish a release date, max-output limit, or deprecation commitments.

Read the AI Wiki article
Official sources (2)
MetaActiveOpen weights

Llama 4 Maverick

llama-4-maverick
Context
1M
Max output
Not published
Input / output
Varies

Capabilities

mixture of expertsmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 400B total / 17B active

License: Llama 4 Community License

Caveat: Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers can impose different limits.

Read the AI Wiki article
Official sources (2)
MetaActiveOpen weights

Llama 4 Scout

llama-4-scout
Context
10M
Max output
Not published
Input / output
Varies

Capabilities

mixture of expertsmultilingualself-hostingfine-tuning

Input: text, image

Output: text

Parameters: 109B total / 17B active

License: Llama 4 Community License

Caveat: Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers may expose a smaller context window.

Read the AI Wiki article
Official sources (2)
MistralActiveAPI + open weights

Mistral Large 3

mistral-large-2512

Aliases: mistral-large-latest

Context
256K
Max output
Not published
Input / output
$0.50 / $1.50

Capabilities

reasoningfunction callingstructured outputsmultilingualself-hosting

Input: text, image

Output: text

Parameters: 675B total / 41B active

License: Apache 2.0

Pricing note: First-party Mistral API price. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.

Caveat: mistral-large-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.

Read the AI Wiki article
Official sources (3)
MistralActiveAPI + open weights

Mistral Medium 3.5

mistral-medium-3-5

Aliases: mistral-medium-3, mistral-medium-latest

Context
256K
Max output
Not published
Input / output
$1.50 / $7.50

Capabilities

reasoningfunction callingstructured outputscodingself-hosting

Input: text, image

Output: text

License: Modified MIT (per the model card)

Pricing note: First-party Mistral API price. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.

Caveat: Mistral publishes dateless IDs only (mistral-medium-3-5 plus the mistral-medium-3 and mistral-medium-latest aliases); the release date and license wording come from the model card alone.

Read the AI Wiki article
Official sources (2)
MistralActiveAPI + open weights

Mistral Small 4

mistral-small-2603

Aliases: mistral-small-latest

Context
256K
Max output
Not published
Input / output
$0.15 / $0.60

Capabilities

reasoningfunction callingstructured outputscodingself-hosting

Input: text, image

Output: text

Parameters: 119B total / 6.5B active

License: Apache 2.0

Pricing note: First-party Mistral API price; self-hosting cost is deployment-specific. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.

Caveat: mistral-small-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.

Read the AI Wiki article
Official sources (2)
OpenAIRetired

GPT-4 32K

gpt-4-32k

Retired

Jun 6, 2025

Past retirement date

Recommended replacement

GPT-4ogpt-4o

OpenAI pointed gpt-4-32k users to gpt-4o, which is itself deprecated with an October 23, 2026 shutdown; new work belongs on the GPT-5.6 family.

Context
32.8K
Max output
Not published
Input / output
Varies

Capabilities

Input: text

Output: text

Caveat: Retired June 6, 2025. Historical pricing is not restated because OpenAI no longer publishes it.

Read the AI Wiki article
Official sources (1)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

OpenAIRetired

OpenAI o1-mini

o1-mini

Retired

Oct 27, 2025

Past retirement date

Announced Apr 28, 2025

Recommended replacement

OpenAI o4-minio4-mini

OpenAI's documented replacement at shutdown was o4-mini, itself now deprecated with an October 23, 2026 shutdown; new work belongs on the GPT-5.6 family.

Context
Not published
Max output
Not published
Input / output
Varies

Capabilities

reasoning

Input: text

Output: text

Caveat: Retired October 27, 2025; requests to this model now error. Historical specifications are not restated here.

Read the AI Wiki article
Official sources (1)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicRetired

Claude Opus 4

claude-opus-4-20250514

Aliases: claude-opus-4

Retired

Jun 15, 2026

Past retirement date

Announced Apr 14, 2026

Recommended replacement

Claude Opus 4.8claude-opus-4-8

Anthropic's documented replacement is Claude Opus 4.8; requests to the retired ID now return an error on Anthropic-operated platforms.

Context
200K
Max output
32K
Input / output
$15 / $75

Capabilities

extended thinkingvisiontool useprompt cachingbatch processing

Input: text, image

Output: text

Pricing note: Anthropic still publishes this rate, labeled retired except on Google Cloud.

Caveat: Retired June 15, 2026 on Anthropic-operated platforms; still listed for Google Cloud. Researchers can request access through Anthropic's External Researcher Access Program.

Read the AI Wiki article
Official sources (3)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

AnthropicRetired

Claude Mythos Preview

claude-mythos-preview

Retired

Jul 21, 2026

Past retirement date

Recommended replacement

Claude Mythos 5claude-mythos-5

Anthropic's documented migration target is Claude Mythos 5, which requires Project Glasswing approval.

Context
1M
Max output
Not published
Input / output
Varies

Capabilities

Input: text, image

Output: text

Caveat: Was an invitation-only research preview for defensive cybersecurity work under Project Glasswing; retired July 21, 2026 and delisted from the pricing page.

Read the AI Wiki article
Official sources (2)

Lifecycle dates above are transcribed from the provider's official notice. Follow the source for last-minute changes.

Methodology

How to read this tracker

Deprecated means a provider has published a replacement and shutdown milestone. Previewmeans the endpoint is usable but may change on a shorter lifecycle. An availability commitment such as Anthropic's "not sooner than" date is shown separately and is not treated as a retirement announcement.

Prices are standard paid text rates in US dollars per one million tokens. Cached reads and Batch API prices appear in the underlying registry when published, while the cards emphasize the comparable standard input/output pair. Long-context tiers and tool charges are called out but not blended into that number.

Model facts are versioned in the site source so price changes, alias changes, and lifecycle updates can be reviewed rather than silently rewriting history. Every record carries its own verification date and first-party links.

Coverage boundaries

What this does not imply

  • This is a curated set of major general-purpose models, not every image, audio, embedding, or experimental endpoint.
  • API lifecycle dates do not automatically apply to ChatGPT, claude.ai, Gemini apps, or third-party cloud catalogs.
  • "Open-weight" describes downloadable weights; it does not guarantee an open-source license or identical hosted limits.
  • A recommended replacement still needs workload-specific evaluation for quality, safety, latency, tools, and cost.

Source index

Provider pricing and lifecycle notices

These are the core documents used across multiple records. Each model card also exposes its specific model card and release sources.

FAQ

Model lifecycle questions

Are all model retirement dates exact?+

No. OpenAI and Anthropic publish confirmed retirement dates for deprecated API models. Google describes dates in its Gemini deprecation table as the earliest possible shutdown dates and says it will communicate the exact date later. The tracker labels that distinction on every affected record.

Does an API model retirement affect ChatGPT, Claude, or Gemini apps?+

Not necessarily. This tracker follows developer model IDs and first-party API lifecycle notices. Consumer apps can switch their underlying models on a separate schedule, and partner platforms such as Amazon Bedrock or Google Cloud can publish different retirement dates.

What happens when an open-weight model is retired?+

Downloaded weights do not disappear. A hosting provider can retire its endpoint, alias, or managed deployment, but self-hosted checkpoints remain available subject to their license. This is why open-weight records show model facts separately from provider-specific hosting cost.

Can I compare the prices as a complete cost estimate?+

The displayed figures are sourced standard text input and output prices per one million tokens. They are a baseline, not a quote: long-context uplifts, tools, regional processing, priority tiers, cache writes, images, audio, and partner margins can change the total.

Planning an API migration?

Export the lifecycle dates above, inventory every hard-coded model ID and alias, and test the replacement against representative production inputs well before the deadline. For related background, browse large language models, OpenAI, Anthropic, and Google DeepMind.

Machine-readable outputs: JSON feed (CC BY 4.0) and a subscribable calendar (.ics) of confirmed retirement dates.