- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $10 / $50
- Released
- Jun 9, 2026
- Knowledge cutoff
- January 2026
Base Claude API price; regional and partner-platform charges can differ.
US export controls forced a global suspension of Fable 5 and Mythos 5 from June 12 to June 30, 2026; access was restored on July 1, 2026. Availability and data-retention eligibility can vary.
Available until at least Jun 9, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
Anthropic
claude-haiku-4-5-20251001 Active- Context
- 200K
- Max output
- 64K
- Distribution
- Hosted API
- Input / output price
- $1 / $5
- Released
- Oct 15, 2025
- Knowledge cutoff
- February 2025
Base Claude API price; regional and partner-platform charges can differ.
The dated API ID is pinned. The shorter claude-haiku-4-5 value is a convenience alias.
Available until at least Oct 15, 2026
text inputimage inputpdf inputextended thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $10 / $50
- Released
- Jun 9, 2026
- Knowledge cutoff
- January 2026
Anthropic lists the same specs and pricing as Claude Fable 5.
Not generally available: offered in limited availability to approved Project Glasswing customers for defensive cybersecurity work. Also covered by the June 12 to June 30, 2026 export-control suspension; access was restored July 1, 2026.
No retirement announced
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $25
- Released
- May 28, 2026
- Knowledge cutoff
- January 2026
Base Claude API price; regional and partner-platform charges can differ.
Available until at least May 28, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2 / $10
- Released
- Jun 30, 2026
- Knowledge cutoff
- January 2026
$2/$10 through Aug 31, 2026, then $3/$15
Introductory pricing through August 31, 2026; Anthropic lists $3 input / $15 output afterward.
Available until at least Jun 30, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
DeepSeek
deepseek-v4-flash Active- Context
- 1M
- Max output
- 384K
- Distribution
- API and downloadable weights
- Input / output price
- $0.14 / $0.28
- Released
- Apr 24, 2026
- Knowledge cutoff
- Not published
Cache-miss input price; cache hits bill at $0.0028 per 1M. One price covers thinking and non-thinking modes.
384K is the published maximum output; DeepSeek does not list a default output limit, knowledge cutoff, or deprecation commitments.
No retirement announced
text inputreasoningself-hosting
- Context
- 1M
- Max output
- 384K
- Distribution
- API and downloadable weights
- Input / output price
- $0.43 / $0.87
- Released
- Apr 24, 2026
- Knowledge cutoff
- Not published
Cache-miss input price; cache hits bill at $0.003625 per 1M. One price covers thinking and non-thinking modes.
384K is the published maximum output; DeepSeek does not list a default output limit, knowledge cutoff, or deprecation commitments.
No retirement announced
text inputreasoningself-hosting
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $1.50 / $9
- Released
- May 19, 2026
- Knowledge cutoff
- Not published
Paid-tier text/image/video input price; feature charges are separate.
gemini-flash-latest is a rolling alias and can change target. Use the stable model ID for production pinning.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Google
gemini-3.5-flash-lite Active- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.30 / $2.50
- Released
- Jul 21, 2026
- Knowledge cutoff
- Not published
Paid-tier text/image/video input price; feature charges are separate. Context caching is paid-tier only (plus $1.00 per 1M tokens per hour of cache storage).
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $1.50 / $7.50
- Released
- Jul 21, 2026
- Knowledge cutoff
- Not published
Paid-tier text/image/video input price; feature charges are separate.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- Varies by host
- Released
- Apr 2, 2026
- Knowledge cutoff
- Not published
Open-weight hosting has no universal price or output cap. The first-party Gemini API deployment can have separate service limits.
No retirement announced
text inputimage inputreasoningfunction callingcoding
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2.50 / $15
- Released
- Not published
- Knowledge cutoff
- August 31, 2025
Higher rates above 272K input tokens ($5 input / $22.50 output).
Positioned by OpenAI as the affordable coding and professional-work model. No published launch date; the dated snapshot id suggests March 5, 2026.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 400K
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $0.75 / $4.50
- Released
- Not published
- Knowledge cutoff
- August 31, 2025
No published launch date; the dated snapshot id suggests March 17, 2026.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 400K
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $0.20 / $1.25
- Released
- Not published
- Knowledge cutoff
- August 31, 2025
No published launch date; the dated snapshot id suggests March 17, 2026.
No retirement announced
text inputimage inputfunction callingstructured outputsstreaming
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $30
- Released
- Not published
- Knowledge cutoff
- December 1, 2025
Higher rates above 272K input tokens ($10 input / $45 output).
Still sold and priced, but no longer shown on OpenAI's models index. OpenAI has not published a launch date; the dated snapshot id suggests April 23, 2026.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $1 / $6
- Released
- Jul 9, 2026
- Knowledge cutoff
- February 16, 2026
Higher rates above 272K input tokens ($2 input / $9 output).
Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $30
- Released
- Jul 9, 2026
- Knowledge cutoff
- February 16, 2026
Higher rates above 272K input tokens ($10 input / $45 output).
Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.
The gpt-5.6 alias currently resolves to Sol; aliases are mutable, so pin gpt-5.6-sol when reproducibility matters.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2.50 / $15
- Released
- Jul 9, 2026
- Knowledge cutoff
- February 16, 2026
Higher rates above 272K input tokens ($5 input / $22.50 output).
Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1M
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $1.25 / $2.50
- Released
- Not published
- Knowledge cutoff
- Not published
Higher rates above 200K input tokens ($2.50 input / $5 output).
grok-latest currently resolves to Grok 4.3, not Grok 4.5. xAI does not publish a release date, max-output limit, or deprecation commitments.
No retirement announced
text inputimage input
- Context
- 500K
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $2 / $6
- Released
- Jul 16, 2026
- Knowledge cutoff
- February 1, 2026
Higher rates above 200K input tokens ($4 input / $12 output).
xAI does not publish max-output limits or deprecation commitments; capability flags are listed on the official model page rather than restated here.
No retirement announced
text inputimage input
- Context
- 256K
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $1 / $2
- Released
- Not published
- Knowledge cutoff
- Not published
Higher rates above 200K input tokens ($2 input / $4 output).
Coding-focused model; legacy grok-code-fast IDs resolve here per xAI's docs. xAI does not publish a release date, max-output limit, or deprecation commitments.
No retirement announced
text inputimage inputcoding
Mistral AI
mistral-large-2512 Active- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- $0.50 / $1.50
- Released
- Dec 2, 2025
- Knowledge cutoff
- Not published
First-party Mistral API price. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.
mistral-large-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
Mistral AI
mistral-medium-3-5 Active- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- $1.50 / $7.50
- Released
- Apr 28, 2026
- Knowledge cutoff
- Not published
First-party Mistral API price. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.
Mistral publishes dateless IDs only (mistral-medium-3-5 plus the mistral-medium-3 and mistral-medium-latest aliases); the release date and license wording come from the model card alone.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
Mistral AI
mistral-small-2603 Active- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- $0.15 / $0.60
- Released
- Mar 16, 2026
- Knowledge cutoff
- Not published
First-party Mistral API price; self-hosting cost is deployment-specific. Cached-input and batch figures are derived from Mistral's published 90 percent cache discount and 50 percent batch discount.
mistral-small-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 131.1K
- Max output
- 131.1K
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Aug 5, 2025
- Knowledge cutoff
- Not published
Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.
No retirement announced
text inputreasoningfunction callingstructured outputs
- Context
- 131.1K
- Max output
- 131.1K
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Aug 5, 2025
- Knowledge cutoff
- Not published
Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.
No retirement announced
text inputreasoningfunction callingstructured outputs
- Context
- 1M
- Max output
- Not published
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Apr 5, 2025
- Knowledge cutoff
- Not published
Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers can impose different limits.
No retirement announced
text inputimage inputmixture of expertsmultilingualself-hosting
- Context
- 10M
- Max output
- Not published
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Apr 5, 2025
- Knowledge cutoff
- Not published
Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers may expose a smaller context window.
No retirement announced
text inputimage inputmixture of expertsmultilingualself-hosting
Google
gemini-3.1-pro-preview Preview- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $2 / $12
- Released
- Feb 19, 2026
- Knowledge cutoff
- Not published
Higher rates above 200K input tokens ($4 input / $18 output).
Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.
Preview models can change or retire on shorter notice. gemini-pro-latest is a rolling alias, not a pinned release.
No retirement date published
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs