- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $10 / $50
- Released
- Jun 9, 2026
- Knowledge cutoff
- January 2026
Base Claude API price; regional and partner-platform charges can differ.
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends Claude Fable 5.1. US export controls forced a global suspension of Fable 5 and Mythos 5 from June 12 to June 30, 2026; access was restored on July 1, 2026.
Available until at least Jun 9, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
Anthropic
claude-fable-5-1 Active- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $10 / $50
- Released
- Sep 1, 2026
- Knowledge cutoff
- June 2026
Base Claude API price. Cache reads cost 0.025x the base input price on this model. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.
Anthropic requires 30-day data retention for this model; it is not available under zero data retention unless Anthropic expressly authorizes it. Forced tool use (tool_choice any or tool) returns a 400 error.
Available until at least Sep 1, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
Anthropic
claude-haiku-4-5-20251001 Active- Context
- 200K
- Max output
- 64K
- Distribution
- Hosted API
- Input / output price
- $1 / $5
- Released
- Oct 15, 2025
- Knowledge cutoff
- February 2025
Base Claude API price; regional and partner-platform charges can differ.
The dated API ID is pinned; claude-haiku-4-5 is a convenience alias. Anthropic's lifecycle table lists it as active with a not-sooner-than date of October 15, 2026, and its model page labels it legacy and recommends Claude Haiku 5.5.
Available until at least Oct 15, 2026
text inputimage inputpdf inputextended thinkingvisiontool use
Anthropic
claude-haiku-5-5 Active- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $0.10 / $0.50
- Released
- Oct 7, 2026
- Knowledge cutoff
- June 2026
Higher rates above 100K input tokens ($0.50 input / $2.50 output).
Base Claude API price for prompts up to 100,000 tokens. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.
Unlike other current Claude models, Haiku 5.5 is priced by prompt length: prompts over 100,000 tokens cost $0.50 input and $2.50 output per million tokens. Manual extended thinking (budget_tokens) returns a 400 error.
Available until at least Oct 7, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $10 / $50
- Released
- Jun 9, 2026
- Knowledge cutoff
- January 2026
Anthropic lists the same specs and pricing as Claude Fable 5.
Not generally available: Anthropic offers it only to organizations verified through its verification programs, such as the Cyber Verification Program, which Anthropic merged with Project Glasswing into one expanded program on October 6, 2026; Claude Mythos 5.1 is now the current Mythos model. Also covered by the June 12 to June 30, 2026 export-control suspension; access was restored July 1, 2026.
Available until at least Jun 9, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
Anthropic
claude-mythos-5-1 Active- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $10 / $50
- Released
- Sep 1, 2026
- Knowledge cutoff
- June 2026
Anthropic lists the same specifications and pricing as Claude Fable 5.1, including cache reads at 0.025x the base input price.
Not generally available: Anthropic offers it only to organizations verified through its verification programs, such as the Cyber Verification Program (whose three access tiers, announced October 6, 2026, each include it) and the Life Sciences Verification Program; at launch it was limited to a set of US organizations. It is the same model as Claude Fable 5.1 with different safeguards, and it carries 30-day data retention unless Anthropic expressly authorizes zero data retention.
Available until at least Sep 1, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
Anthropic
claude-opus-4-5-20251101 Active- Context
- 200K
- Max output
- 64K
- Distribution
- Hosted API
- Input / output price
- $5 / $25
- Released
- Nov 24, 2025
- Knowledge cutoff
- May 2025
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.
Available until at least Nov 24, 2026
text inputimage inputpdf inputextended thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $25
- Released
- Feb 5, 2026
- Knowledge cutoff
- May 2025
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.
Available until at least Feb 5, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $25
- Released
- Apr 16, 2026
- Knowledge cutoff
- January 2026
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models. Fast mode was deprecated June 25, 2026, with removal on July 24, 2026.
Available until at least Apr 16, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $25
- Released
- May 28, 2026
- Knowledge cutoff
- January 2026
Base Claude API price; regional and partner-platform charges can differ.
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends Claude Opus 5.5.
Available until at least May 28, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $25
- Released
- Jul 24, 2026
- Knowledge cutoff
- May 2026
Base Claude API price, the same as Claude Opus 4.8. Fast mode (research preview, Claude API only) costs $10 input and $50 output per million tokens. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends migrating to Claude Opus 5.5. Disabling thinking is allowed only at effort high or below.
Available until at least Jul 24, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $4 / $20
- Released
- Sep 22, 2026
- Knowledge cutoff
- June 2026
Base Claude API price. Cache reads cost 0.05x the base input price on this model. Fast mode (research preview, Claude API only) costs $8 input and $40 output per million tokens. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.
Adaptive thinking is always on and cannot be turned off, and forced tool use (tool_choice any or tool) returns a 400 error. The Message Batches API allows up to 300K output tokens with the output-300k-2026-03-24 beta header.
Available until at least Sep 22, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
Anthropic
claude-sonnet-4-6 Active- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $3 / $15
- Released
- Feb 17, 2026
- Knowledge cutoff
- August 2025
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; the models overview groups it under legacy models.
Available until at least Feb 17, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2 / $10
- Released
- Jun 30, 2026
- Knowledge cutoff
- January 2026
These are standard prices. The previously announced September 1 increase was cancelled.
Anthropic's lifecycle table lists this model as active with a not-sooner-than commitment; its model page labels it legacy and recommends Claude Sonnet 5.5.
Available until at least Jun 30, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
Anthropic
claude-sonnet-5-5 Active- Context
- 1M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2 / $10
- Released
- Sep 28, 2026
- Knowledge cutoff
- June 2026
Base Claude API price. Cache reads were $0.20 per million tokens at launch on September 28, 2026 and were cut to $0.10 (0.05x the base input price) on October 7, 2026; all other rates are unchanged since launch. US-only inference (inference_geo) costs 1.1x, and partner-platform prices can differ.
Forced tool use (tool_choice any or tool) returns a 400 error, and its thinking blocks work only in the account that produced them or a linked account. The Message Batches API allows up to 300K output tokens with the output-300k-2026-03-24 beta header.
Available until at least Sep 28, 2027
text inputimage inputpdf inputadaptive thinkingvisiontool use
- Context
- 1M
- Max output
- 384K
- Distribution
- API and downloadable weights
- Input / output price
- $1.32 / $3.96
- Released
- Apr 24, 2026
- Knowledge cutoff
- Not published
Peak rates used for planning. Off-peak input/cache/output: $0.66/$0.022/$1.98 (half the peak rates). Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays; weekends and Chinese public holidays are off-peak all day.
Served by the DeepSeek-V4-Pro-0813 build since August 13, 2026; image input is not supported. DeepSeek's September 10 news post still says this name routes to V4.1-Flash from September 14, but its change log says V4-Pro API service continues after that date with billing unchanged until further notice, and the pricing page still lists it.
No retirement announced
text inputreasoningtool usestructured outputs
- Context
- 1M
- Max output
- 384K
- Distribution
- API and downloadable weights
- Input / output price
- $0.30 / $1.20
- Released
- Sep 10, 2026
- Knowledge cutoff
- Not published
Peak rates used for planning. Off-peak input/cache/output: $0.15/$0.003/$0.60 (half the peak rates). Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays; weekends and Chinese public holidays are off-peak all day.
API model ID is deepseek-flash; the retired names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, served by this model and billed at these rates. Generic dense-transformer memory formulas do not model this architecture accurately.
No retirement announced
text inputimage inputreasoningtool usestructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.30 / $2.50
- Released
- Jun 17, 2025
- Knowledge cutoff
- Not published
Google's September 18, 2026 release notes limit access to the 2.5 models to users who have actively used them in the past. Google says they are not deprecated and lists no shutdown date, and it directs new projects to Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Google
gemini-2.5-flash-lite Active- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.10 / $0.40
- Released
- Jul 22, 2025
- Knowledge cutoff
- Not published
Google's September 18, 2026 release notes limit access to the 2.5 models to users who have actively used them in the past. Google says they are not deprecated and lists no shutdown date, and it directs new projects to Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $1.25 / $10
- Released
- Jun 17, 2025
- Knowledge cutoff
- Not published
Higher rates above 200K input tokens ($2.50 input / $15 output).
Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.
Google's September 18, 2026 release notes limit access to the 2.5 models to users who have actively used them in the past. Google says they are not deprecated and lists no shutdown date, and it directs new projects to Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $1.50 / $9
- Released
- May 19, 2026
- Knowledge cutoff
- Not published
Paid-tier text/image/video input price; feature charges are separate.
Google's models page describes this as its legacy Flash model, but it is still listed as stable and the deprecation table has no shutdown date. Google's May 19, 2026 release notes made it the model behind gemini-flash-latest; latest aliases can be hot-swapped on new releases, so pin the stable ID.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Google
gemini-3.5-flash-lite Active- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.30 / $2.50
- Released
- Jul 21, 2026
- Knowledge cutoff
- Not published
Paid-tier text/image/video input price; feature charges are separate. Context caching is paid-tier only (plus $1.00 per 1M tokens per hour of cache storage).
The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.75 / $3.75
- Released
- Jul 21, 2026
- Knowledge cutoff
- Not published
$0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50
Google's Gemini 3.7 Flash guide says the 3.7 introductory rate also applies to 3.6 Flash. The first dated listing is the 3.7 Flash model card published August 13, 2026; Google states no separate start date. Launch list price was $1.50 input and $7.50 output. Cache storage and grounding are billed separately.
The current Gemini API deprecation table has no shutdown date for this model. Model and service limits remain subject to the linked documentation.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.75 / $3.75
- Released
- Aug 13, 2026
- Knowledge cutoff
- March 2026
$0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50
Introductory paid-tier Standard rates through December 31, 2026; the output price includes thinking tokens. Cache storage ($0.50 per 1M tokens per hour) and grounding are billed separately.
Google qualifies the March 2026 knowledge cutoff: some domains may be limited to January 2025. Google says 3.7 Flash remains fully supported for efficiency-first workloads alongside the newer 3.8 Flash.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $0.75 / $3.75
- Released
- Sep 2, 2026
- Knowledge cutoff
- March 2026
$0.75/$3.75 through Dec 31, 2026, then $1.50/$7.50
Introductory paid-tier Standard rates announced at launch and valid through December 31, 2026; the output price includes thinking tokens. Cache storage ($0.50 per 1M tokens per hour) and grounding are billed separately.
Google qualifies the March 2026 knowledge cutoff: some domains may be limited to January 2025. Google warns the model may use more tokens on complex tasks, especially at higher effort levels.
No retirement announced
text inputimage inputaudio inputvideo inputpdf inputthinkingfunction callingstructured outputs
- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- Varies by host
- Released
- Apr 2, 2026
- Knowledge cutoff
- Not published
Open-weight hosting has no universal price or output cap. The first-party Gemini API deployment can have separate service limits.
No retirement announced
text inputimage inputreasoningfunction callingcoding
- Context
- 128K
- Max output
- 16.4K
- Distribution
- Hosted API
- Input / output price
- $2.50 / $10
- Released
- May 13, 2024
- Knowledge cutoff
- October 1, 2023
Rates for the gpt-4o alias as listed on October 8, 2026; the pricing page does not say when they took effect. The deprecated gpt-4o-2024-05-13 snapshot bills $5 input and $15 output until it shuts down.
Only the older gpt-4o-2024-05-13 snapshot is deprecated, with an October 23, 2026 shutdown. The gpt-4o alias points to gpt-4o-2024-08-06 and is not on OpenAI's deprecation list.
No retirement announced
text inputimage inputfunction callingstructured outputsstreaming
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2.50 / $15
- Released
- Not published
- Knowledge cutoff
- August 31, 2025
Higher rates above 272K input tokens ($5 input / $22.50 output).
Positioned by OpenAI as the affordable coding and professional-work model. No published launch date; the dated snapshot id suggests March 5, 2026.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 400K
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $0.75 / $4.50
- Released
- Not published
- Knowledge cutoff
- August 31, 2025
No published launch date; the dated snapshot id suggests March 17, 2026.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $5 / $30
- Released
- Not published
- Knowledge cutoff
- December 1, 2025
Higher rates above 272K input tokens ($10 input / $45 output).
Still sold and priced, but no longer shown on OpenAI's models index. OpenAI has not published a launch date; the dated snapshot id suggests April 23, 2026.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $0.20 / $1.20
- Released
- Jul 9, 2026
- Knowledge cutoff
- February 16, 2026
Higher rates above 272K input tokens ($0.40 input / $1.80 output).
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $4 / $20
- Released
- Jul 9, 2026
- Knowledge cutoff
- February 16, 2026
Higher rates above 272K input tokens ($8 input / $30 output).
Promotional rates are available at least through November 21, 2026. Later rates have not been committed; projections hold these rates constant.
The gpt-5.6 alias currently resolves to Sol; aliases are mutable, so pin gpt-5.6-sol when reproducibility matters.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2 / $12
- Released
- Jul 9, 2026
- Knowledge cutoff
- February 16, 2026
Higher rates above 272K input tokens ($4 input / $18 output).
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $10 / $50
- Released
- Sep 3, 2026
- Knowledge cutoff
- April 30, 2026
Higher rates above 272K input tokens ($20 input / $75 output).
Flex is priced like Batch and Fast mode costs 2x Standard. Since September 29, 2026 an Ultrafast tier (service_tier ultrafast, Responses API) costs $60 input, $6 cached input, $75 cache writes and $300 output per 1M tokens, or $120, $12, $150 and $450 above 272K input tokens.
Does not accept the none reasoning effort, custom temperature or top_p values, or log probabilities, and tool calling requires the Responses API (Chat Completions works without tools).
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $0.10 / $0.50
- Released
- Sep 22, 2026
- Knowledge cutoff
- May 18, 2026
Higher rates above 272K input tokens ($0.20 input / $0.75 output).
Flex is priced like Batch, Fast mode costs 2x Standard, and regional processing adds 10 percent where available.
On September 25, 2026 OpenAI fixed an image-encoding bug that had degraded image understanding, so rerun image evaluations made before then. Chat Completions supports function calling only with reasoning effort none; use the Responses API for tools with reasoning.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2 / $10
- Released
- Sep 22, 2026
- Knowledge cutoff
- April 20, 2026
Higher rates above 272K input tokens ($4 input / $15 output).
Flex is priced like Batch, Fast mode costs 2x Standard, and regional processing adds 10 percent where available.
OpenAI's docs now point to GPT-6.1 Sol as the newer Sol model, though gpt-6-sol has no announced deprecation. On September 25, 2026 OpenAI fixed an image-encoding bug that had degraded image understanding, so rerun image evaluations made before then.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- 128K
- Distribution
- Hosted API
- Input / output price
- $2 / $10
- Released
- Sep 29, 2026
- Knowledge cutoff
- April 30, 2026
Higher rates above 272K input tokens ($4 input / $15 output).
Cached input is billed at 5 percent of the input rate, half of GPT-6 Sol's cached price. Flex is priced like Batch, Fast mode costs 2x Standard, and regional processing adds 10 percent where available.
Does not accept the none or minimal reasoning efforts, and tool calling requires the Responses API (Chat Completions works without tools).
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1M
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $1.25 / $2.50
- Released
- Not published
- Knowledge cutoff
- Not published
Higher rates at or above 200K input tokens ($2.50 input / $5 output).
The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 500K
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $2 / $6
- Released
- Jul 16, 2026
- Knowledge cutoff
- February 1, 2026
Higher rates at or above 200K input tokens ($4 input / $12 output).
The official model page lists no maximum output limit or committed retirement date. Alias mappings can change; use the explicit model ID.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 500K
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $2 / $6
- Released
- Aug 12, 2026
- Knowledge cutoff
- February 1, 2026
Higher rates at or above 200K input tokens ($4 input / $12 output).
No batch discount: the Batch API is not supported for this model. Requests to the US regional endpoint (https://us.api.x.ai/v1) bill at 1.1x these rates, Priority Processing bills at 2x, and server-side tool calls are billed separately.
xAI states that this model has no text output limit and no Batch API support, and xAI publishes no retirement date for it. Its overview page now redirects to Grok 4.7 and the models page recommends Grok 4.7, but grok-4.6 keeps its model page, its pricing row and its place on the US regional endpoint.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 500K
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $2 / $6
- Released
- Sep 21, 2026
- Knowledge cutoff
- May 2026
Higher rates at or above 200K input tokens ($4 input / $12 output).
No batch discount: the Batch API is not supported for this model. Requests to the US regional endpoint (https://us.api.x.ai/v1) bill at 1.1x these rates, Priority Processing bills at 2x, and server-side tool calls are billed separately.
xAI states that this model has no text output limit and no Batch API support, and xAI publishes no retirement date for it. Grok 4.7 Fast (the same model at 2x standard rates, or 1.5x for long-context requests) is sold only in Cursor and Grok Build, not on the public xAI API.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 256K
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $1 / $2
- Released
- May 29, 2026
- Knowledge cutoff
- Not published
Higher rates at or above 200K input tokens ($2 input / $4 output).
xAI announced grok-build-0.1 on its API in public beta on May 29, 2026 (its May release notes call it early access). The official model page lists no maximum output limit or committed retirement date, and alias mappings can change, so use the explicit model ID.
No retirement announced
text inputimage inputcodingreasoningfunction calling
Mistral AI
mistral-large-2512 Active- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- $0.50 / $1.50
- Released
- Dec 2, 2025
- Knowledge cutoff
- Not published
First-party Mistral API price. Cached-input and batch rates are published directly on Mistral's pricing page (Standard and Batch views).
mistral-large-latest is a rolling alias that still points to Mistral Large 3 (mistral-large-2512), not to the Mistral Large 4 public preview. The published model card does not specify a separate maximum-output limit.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
Mistral AI
mistral-medium-3-5 Active- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- $1.50 / $7.50
- Released
- Apr 28, 2026
- Knowledge cutoff
- Not published
First-party Mistral API price. Cached-input and batch rates are published directly on Mistral's pricing page (Standard and Batch views).
Mistral publishes dateless IDs only (mistral-medium-3-5 plus the mistral-medium-3 and mistral-medium-latest aliases); the release date and license wording come from the model card alone.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
Mistral AI
mistral-small-2603 Active- Context
- 256K
- Max output
- Not published
- Distribution
- API and downloadable weights
- Input / output price
- $0.15 / $0.60
- Released
- Mar 16, 2026
- Knowledge cutoff
- Not published
First-party Mistral API price; self-hosting cost is deployment-specific. Cached-input and batch rates are published directly on Mistral's pricing page (Standard and Batch views).
mistral-small-latest is a rolling alias. The published model card does not specify a separate maximum-output limit.
No retirement announced
text inputimage inputreasoningfunction callingstructured outputs
- Context
- 1.05M
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $1.25 / $4.25
- Released
- Sep 2, 2026
- Knowledge cutoff
- Not published
Standard tier rates as observed on October 8, 2026 (Meta's pricing page does not date them); prompts and completions are not used to train Meta models. No long-context premium; built-in web search grounding adds $2.50 per 1,000 queries. The muse-spark-1.3-contributor variant bills $0.10/$0.002/$0.20 in exchange for training use.
Meta lists audio input but says audio understanding in Muse Spark 1.3 is not fully supported and may be degraded (it points audio work to Muse Spark 1.2 or Muse Voice Transcribe). Meta does not publish a maximum output limit, knowledge cutoff, or parameter count, and the max reasoning level is Standard-tier only.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputreasoningfunction callingstructured outputs
Meta
muse-spark-1.3-contributor Active- Context
- 1.05M
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $0.10 / $0.20
- Released
- Not published
- Knowledge cutoff
- Not published
Contributor tier rates as observed on October 8, 2026 (Meta's pricing page does not date them): discounted in exchange for permission to use prompts and completions to train future Meta models. The same model on the Standard tier (muse-spark-1.3) costs $1.25/$0.15/$4.25. Built-in web search grounding adds $2.50 per 1,000 queries.
Training-eligible tier of Muse Spark 1.3: Meta may use your prompts and completions to train future Meta models, the max reasoning level is not available, and the rate limit is 100 requests and 3,000,000 tokens per minute. Meta does not date this tier's launch, and it says audio understanding in Muse Spark 1.3 is not fully supported.
No retirement announced
text inputimage inputvideo inputaudio inputpdf inputreasoningfunction callingstructured outputs
- Context
- 131.1K
- Max output
- 131.1K
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Aug 5, 2025
- Knowledge cutoff
- Not published
Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.
No retirement announced
text inputreasoningfunction callingstructured outputs
- Context
- 131.1K
- Max output
- 131.1K
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Aug 5, 2025
- Knowledge cutoff
- Not published
Open weights have no universal token price. Hosting cost, quantization, context limits, and throughput depend on the deployment.
No retirement announced
text inputreasoningfunction callingstructured outputs
- Context
- 1M
- Max output
- Not published
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Apr 5, 2025
- Knowledge cutoff
- Not published
Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers can impose different limits.
No retirement announced
text inputimage inputmixture of expertsmultilingualself-hosting
- Context
- 10M
- Max output
- Not published
- Distribution
- Open weights
- Input / output price
- Varies by host
- Released
- Apr 5, 2025
- Knowledge cutoff
- Not published
Llama is open-weight under Meta's custom community license, not OSI open source. Hosted providers may expose a smaller context window.
No retirement announced
text inputimage inputmixture of expertsmultilingualself-hosting
Google
gemini-3.1-pro-preview Preview- Context
- 1.05M
- Max output
- 65.5K
- Distribution
- Hosted API
- Input / output price
- $2 / $12
- Released
- Feb 19, 2026
- Knowledge cutoff
- Not published
Higher rates above 200K input tokens ($4 input / $18 output).
Paid-tier price for prompts up to 200K tokens. Longer prompts use higher input and output rates.
Preview models can change or retire on shorter notice. gemini-pro-latest is a rolling alias, not a pinned release.
No retirement date published
text inputimage inputvideo inputaudio inputpdf inputthinkingfunction callingstructured outputs
Mistral AI
mistral-large-4 Preview- Context
- 1M
- Max output
- Not published
- Distribution
- Hosted API
- Input / output price
- $1.36 / $4.18
- Released
- Oct 6, 2026
- Knowledge cutoff
- Not published
Standard (original) Mistral API rates; batch cached input is $0.07. Mistral's changelog announces launch pricing of 50% off for 2 weeks, shown on the pricing page as a temporary sale ($0.68 input, $0.07 cached input, $2.09 output; batch $0.34/$0.035/$1.045). Mistral publishes no exact end date for the sale, so planning uses the standard rates.
Public Preview: Mistral allows silent updates and does not guarantee general availability, and the model card lists the weights and license as coming soon (Mistral says weights ship by the end of October 2026). Mistral publishes no maximum output limit, and mistral-large-latest still points to Mistral Large 3.
No retirement date published
text inputimage inputreasoningfunction callingstructured outputs