- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $10 / $50
- Released
- Jun 9, 2026
- Knowledge cutoff
- January 2026
Base Claude API price; regional and partner-platform charges can differ.
US export controls forced a global suspension of Fable 5 and Mythos 5 from June 12 to June 30, 2026; access was restored on July 1, 2026. Availability and data-retention eligibility can vary.
Available until at least Jun 9, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
Anthropic
claude-haiku-4-5-20251001 Active- Context
- 200K
- Max output
- 64K
- Distribution
- Hosted API
- Input / output price
- $1 / $5
- Released
- Oct 15, 2025
- Knowledge cutoff
- February 2025
Base Claude API price; regional and partner-platform charges can differ.
The dated API ID is pinned. The shorter claude-haiku-4-5 value is a convenience alias.
Available until at least Oct 15, 2026
text inputimage inputpdf inputextended thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $10 / $50
- Released
- Jun 9, 2026
- Knowledge cutoff
- January 2026
Anthropic lists the same specs and pricing as Claude Fable 5.
Not generally available: offered in limited availability to approved Project Glasswing customers for defensive cybersecurity work. Also covered by the June 12 to June 30, 2026 export-control suspension; access was restored July 1, 2026.
No retirement announced
text inputimage inputpdf inputadaptive thinkingvisiontool use
Anthropic
claude-opus-4-5-20251101 Active- Context
- 200K
- Max output
- 64K
- Distribution
- Hosted API
- Input / output price
- $5 / $25
- Released
- Nov 24, 2025
- Knowledge cutoff
- May 2025
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.
Available until at least Nov 24, 2026
text inputimage inputpdf inputextended thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $25
- Released
- Feb 5, 2026
- Knowledge cutoff
- May 2025
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.
Available until at least Feb 5, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $25
- Released
- Apr 16, 2026
- Knowledge cutoff
- January 2026
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models. Fast mode was deprecated June 25, 2026, with removal on July 24, 2026.
Available until at least Apr 16, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $25
- Released
- May 28, 2026
- Knowledge cutoff
- January 2026
Base Claude API price; regional and partner-platform charges can differ.
Available until at least May 28, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
Anthropic
claude-sonnet-4-5-20250929 Active- Context
- 200K
- Max output
- 64K
- Distribution
- Hosted API
- Input / output price
- $3 / $15
- Released
- Sep 29, 2025
- Knowledge cutoff
- January 2025
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.
Available until at least Sep 29, 2026
text inputimage inputpdf inputextended thinkingvisiontool use
Anthropic
claude-sonnet-4-6 Active- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $3 / $15
- Released
- Feb 17, 2026
- Knowledge cutoff
- August 2025
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.
Available until at least Feb 17, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2 / $10
- Released
- Jun 30, 2026
- Knowledge cutoff
- January 2026
These are standard prices. The previously announced September 1 increase was cancelled.
Available until at least Jun 30, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 384K
- Distribution
- API and downloadable weights
- Input / output price
- $0.30 / $1.20
- Released
- Sep 10, 2026
- Knowledge cutoff
- Not published
Peak rates used for planning. Off-peak input/cache/output: $0.15/$0.003/$0.60. Peak hours: Monday-Friday 01:00-04:00 and 06:00-10:00 UTC; all other hours are half price.
API model ID is deepseek-flash. Generic dense-transformer memory formulas do not model this architecture accurately.
No retirement announced
text inputimage inputreasoningtool usestructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.30 / $2.50
- Released
- Jun 17, 2025
- Knowledge cutoff
- Not published
The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Google
gemini-2.5-flash-lite Active- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.10 / $0.40
- Released
- Jul 22, 2025
- Knowledge cutoff
- Not published
The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $1.25 / $10
- Released
- Jun 17, 2025
- Knowledge cutoff
- Not published
Higher rates above 200K input tokens ($2.50 input / $15 output).
Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.
The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $1.50 / $9
- Released
- May 19, 2026
- Knowledge cutoff
- Not published
Paid-tier text/image/video input price; feature charges are separate.
The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Google
gemini-3.5-flash-lite Active- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.30 / $2.50
- Released
- Jul 21, 2026
- Knowledge cutoff
- Not published
Paid-tier text/image/video input price; feature charges are separate. Context caching is paid-tier only (plus $1.00 per 1M tokens per hour of cache storage).
The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.75 / $3.75
- Released
- Jul 21, 2026
- Knowledge cutoff
- Not published
$0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50
Rates observed September 11; the pricing page does not state when the discount began. Cache storage and grounding are additional charges.
The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.75 / $3.75
- Released
- Sep 2, 2026
- Knowledge cutoff
- Not published
$0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50
Rates observed September 11; the pricing page does not state when the discount began. Cache storage and grounding are additional charges.
No retirement announced
text inputimage inputaudio inputvideo inputpdf inputthinkingfunction callingstructured outputs
- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- Varies by host
- Released
- Apr 2, 2026
- Knowledge cutoff
- Not published
Open-weight hosting has no universal price or output cap. The first-party Gemini API deployment can have separate service limits.
No retirement announced
text inputimage inputreasoningfunction callingcoding
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2.50 / $15
- Released
- Not published
- Knowledge cutoff
- August 31, 2025
Higher rates above 272K input tokens ($5 input / $22.50 output).
Positioned by OpenAI as the affordable coding and professional-work model. No published launch date; the dated snapshot id suggests March 5, 2026.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 400K
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $0.75 / $4.50
- Released
- Not published
- Knowledge cutoff
- August 31, 2025
No published launch date; the dated snapshot id suggests March 17, 2026.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 400K
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $0.20 / $1.25
- Released
- Not published
- Knowledge cutoff
- August 31, 2025
No published launch date; the dated snapshot id suggests March 17, 2026.
No retirement announced
text inputimage inputfunction callingstructured outputsstreaming
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $30
- Released
- Not published
- Knowledge cutoff
- December 1, 2025
Higher rates above 272K input tokens ($10 input / $45 output).
Still sold and priced, but no longer shown on OpenAI's models index. OpenAI has not published a launch date; the dated snapshot id suggests April 23, 2026.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $0.20 / $1.20
- Released
- Jul 9, 2026
- Knowledge cutoff
- February 16, 2026
Higher rates above 272K input tokens ($0.40 input / $1.80 output).
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $4 / $20
- Released
- Jul 9, 2026
- Knowledge cutoff
- February 16, 2026
Higher rates above 272K input tokens ($8 input / $30 output).
Promotional rates are available at least through November 21, 2026. Later rates have not been committed; projections hold these rates constant.
The gpt-5.6 alias currently resolves to Sol; aliases are mutable, so pin gpt-5.6-sol when reproducibility matters.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2 / $12
- Released
- Jul 9, 2026
- Knowledge cutoff
- February 16, 2026
Higher rates above 272K input tokens ($4 input / $18 output).
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $10 / $50
- Released
- Not published
- Knowledge cutoff
- April 30, 2026
Higher rates above 272K input tokens ($20 input / $75 output).
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1M
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $1.25 / $2.50
- Released
- Not published
- Knowledge cutoff
- Not published
Higher rates at or above 200K input tokens ($2.50 input / $5 output).
The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.
No retirement announced
text inputimage input
- Context
- 500K
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $2 / $6
- Released
- Jul 16, 2026
- Knowledge cutoff
- February 1, 2026
Higher rates at or above 200K input tokens ($4 input / $12 output).
The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.
No retirement announced
text inputimage input
- Context
- 256K
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $1 / $2
- Released
- Not published
- Knowledge cutoff
- Not published
Higher rates at or above 200K input tokens ($2 input / $4 output).
The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.
No retirement announced
text inputimage inputcoding
Mistral AI
mistral-large-2512 Active- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- $0.50 / $1.50
- Released
- Dec 2, 2025
- Knowledge cutoff
- Not published
First-party Mistral API price. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.
mistral-large-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
Mistral AI
mistral-medium-3-5 Active- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- $1.50 / $7.50
- Released
- Apr 28, 2026
- Knowledge cutoff
- Not published
First-party Mistral API price. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.
Mistral publishes dateless IDs only (mistral-medium-3-5 plus the mistral-medium-3 and mistral-medium-latest aliases); the release date and license wording come from the model card alone.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
Mistral AI
mistral-small-2603 Active- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- $0.15 / $0.60
- Released
- Mar 16, 2026
- Knowledge cutoff
- Not published
First-party Mistral API price; self-hosting cost is deployment-specific. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.
mistral-small-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 131.1K
- Max output
- 131.1K
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Aug 5, 2025
- Knowledge cutoff
- Not published
Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.
No retirement announced
text inputreasoningfunction callingstructured outputs
- Context
- 131.1K
- Max output
- 131.1K
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Aug 5, 2025
- Knowledge cutoff
- Not published
Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.
No retirement announced
text inputreasoningfunction callingstructured outputs
- Context
- 1M
- Max output
- Not published
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Apr 5, 2025
- Knowledge cutoff
- Not published
Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers can impose different limits.
No retirement announced
text inputimage inputmixture of expertsmultilingualself-hosting
- Context
- 10M
- Max output
- Not published
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Apr 5, 2025
- Knowledge cutoff
- Not published
Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers may expose a smaller context window.
No retirement announced
text inputimage inputmixture of expertsmultilingualself-hosting
Google
gemini-3.1-pro-preview Preview- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $2 / $12
- Released
- Feb 19, 2026
- Knowledge cutoff
- Not published
Higher rates above 200K input tokens ($4 input / $18 output).
Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.
Preview models can change or retire on shorter notice. gemini-pro-latest is a rolling alias, not a pinned release.
No retirement date published
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs