LLM API Pricing Comparison
Every figure in this article is a US dollar list rate per 1,000,000 tokens, taken from the provider's own pricing page. The full table was checked on August 1, 2026, and the OpenAI, Anthropic, Google, DeepSeek, and xAI flagship rows were re-checked on September 3, 2026, the day OpenAI released GPT-6 Astra. Those dates matter more than any single number below. In the five weeks between the two checks, OpenAI cut GPT-5.6 Sol's output price by a third and then priced Astra at 2.5 times Sol's promotional rate, Anthropic cancelled a scheduled 50 percent increase on Claude Sonnet 5 that an earlier revision of this page listed for September 1, Google halved Gemini 3.6 Flash through the end of the year, and DeepSeek moved to peak and off-peak billing with higher base rates. Prices in this market change without notice, sometimes twice in a quarter, so treat the tables as a snapshot and re-check the linked source pages before you commit a budget.
The short answer as of September 3: the cheapest capable API is still DeepSeek V4-Flash, but at $0.22 input and $0.66 output off-peak (double that during weekday peak hours), not the $0.14 and $0.28 it carried in July, and its cache-hit input rate of $0.007 remains the deepest cache discount any major provider publishes.[8] The cheapest frontier-class model is DeepSeek V4-Pro at $0.66 and $1.98 off-peak, and both ship as open weights under the MIT license, so the API rate is only one way to buy them.[8][28] Among the largest US labs, OpenAI's GPT-5.6 Luna at $0.20 and $1.20 is the cheapest model that still scores near the top of independent aggregate rankings, and Google's Gemini 3.8 Flash sits just above it at an introductory $0.75 and $3.75.[1][7] The premium tier now has two models at exactly the same list price: GPT-6 Astra and Anthropic's Claude Fable 5.1, both $10 input and $50 output, with GPT-5.6 Sol at $4 and $20 on a promotion that runs at least through November 21, 2026 and Claude Opus 5 at $5 and $25 below them.[1][6][33] Google's Gemini 3.1 Pro undercuts all of those at $2 and $12 for prompts up to 200,000 tokens.[7]
That is where the naive comparison ends and the useful one begins. A per-token rate is now a weak predictor of what a job costs, because two models at similar rates can emit output token counts that differ by two orders of magnitude for the same task. The sections after the tables cover the things that actually move a bill: reasoning effort, cache state, prompt-length thresholds, service tiers, time of day, and where you buy the model.
This is the developer, per-token API pricing hub. For AI subscriptions, consumer plans, and market pricing generally, see the AI pricing concept page.
Scope, units, and how volatile this is
Unless a row says otherwise, every price is:
- a public list rate, not a negotiated enterprise rate;
- in US dollars per 1,000,000 tokens;
- on the standard synchronous tier (batch, flex, and priority tiers are covered separately);
- on the provider's own first-party API, not a reseller or cloud marketplace;
- at the short-context tier, where a provider charges more above a token threshold;
- at the off-peak rate, where a provider charges by time of day.
Where a provider does not publish a distinct cache-hit rate, the tables say "not published" rather than assuming one. Where a provider's current price could not be confirmed on an official page, that provider is left out entirely rather than estimated. Several vendors that appeared in earlier revisions of this article, including xAI's Grok 4.1 Fast, are no longer listed on the vendor's own model page and have been removed for that reason.
Recent and scheduled changes to the rates below:
| Date | Change | Provider |
|---|---|---|
| July 30, 2026 | GPT-5.6 Luna list price cut 80 percent; GPT-5.6 Terra cut 20 percent[2] | OpenAI |
| July 30, 2026 | "Priority processing" renamed "Fast mode"; service_tier: "priority" still accepted[1][30] | OpenAI |
| August 13, 2026 | Gemini 3.7 Flash launched at an introductory $0.75 / $3.75, "half the original 3.6 Flash cost"; the same rate now applies to Gemini 3.6 Flash through December 31, 2026, with $1.50 / $7.50 scheduled from January 1, 2027[7][37] | |
| August 13, 2026 | "Ultrafast mode" announced for GPT-5.6 Sol, up to 14x standard speed, limited preview, no price published[2] | OpenAI |
| August 16, 2026 | Peak and off-peak billing took effect at 16:00 UTC with new base rates; off-peak is half of peak[8][36] | DeepSeek |
| August 21, 2026 | GPT-5.6 Sol cut to a promotional $4 / $0.40 / $20, "at least through November 21, 2026"[2] | OpenAI |
| By September 3, 2026 | Claude Sonnet 5's $2 / $10 introductory rate made permanent; the $3 / $15 increase scheduled for September 1 "will not occur"[6] | Anthropic |
| By September 3, 2026 | Claude Fable 5.1 and Claude Mythos 5.1 listed at $10 / $50 with cache reads at $0.25, a 0.025x multiplier against the usual 0.1x[6] | Anthropic |
| September 2, 2026 | Gemini 3.8 Flash launched "at the same introductory price as 3.7 Flash", $0.75 / $3.75 through December 31, 2026[7][41] | |
| September 3, 2026 | GPT-6 Astra released as gpt-6-astra at $10 / $1 / $12.50 / $50 (input, cached input, cache write, output); Fast mode at 2x[1][2][32] | OpenAI |
What API pricing actually varies on
Nine things move a bill. Only the first is a headline number.
| Axis | What it is | Size of the effect |
|---|---|---|
| Direction of the token | Output is billed above input everywhere | 2x to about 8x across the models below |
| Cache state | Cache miss, cache write, and cache hit are three different prices | Hits are 90 percent off at most providers, 97.5 percent off on Claude Fable 5.1, and about 97 percent off at DeepSeek; writes cost 1.25x to 2x list at Anthropic and on GPT-5.6 and GPT-6 Astra |
| Prompt length | A threshold above which the whole request reprices | OpenAI doubles input above 272,000 tokens; Google and xAI double above 200,000; Anthropic charges flat to 1M |
| Service tier | Batch, flex, standard, fast | Batch and flex are half price; fast and priority run 1.8x to 2.5x |
| Time of day | Peak-hour surcharges | DeepSeek charges double during weekday peak hours (01:00-04:00 and 06:00-10:00 UTC) |
| Point of sale | First-party API, cloud marketplace, third-party host | Together AI charged exactly 4x DeepSeek's own July rate for the same V4-Pro weights; regional endpoints add 10 percent at OpenAI and Anthropic |
| Reasoning effort | How many output tokens the model spends before answering | 68x on a single prompt within one model family |
| Tokenizer | How many tokens your text becomes | About 30 percent more on Claude 4.7 and later |
| Non-token line items | Server-side tools and session runtime | $10 per 1,000 web searches, $0.08 per session-hour, hourly cache storage at Google |
The rest of this article works through those in order of how much money they tend to move.
The cross-provider price table
Standard tier, short-context rate, sorted by output price. Cached input is the cache-hit read rate, not the cache-write rate. DeepSeek appears twice because it now bills by time of day.
| Model | Provider | Input | Cached input | Output |
|---|---|---|---|---|
| GPT-5-nano (deprecated) | OpenAI | $0.05 | $0.005 | $0.40 |
| Qwen-Flash (0-256K) | Alibaba | $0.05 | not published | $0.40 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.01 | $0.40 | |
| Mistral Small 4 | Mistral AI | $0.15 | not published | $0.60 |
| DeepSeek V4-Flash (off-peak) | DeepSeek | $0.22 | $0.007 | $0.66 |
| GLM-4.5-Air | Z.ai | $0.20 | $0.03 | $1.10 |
| GPT-5.6 Luna | OpenAI | $0.20 | $0.02 | $1.20 |
| Qwen-Plus (0-256K) | Alibaba | $0.40 | not published | $1.20 non-thinking, $4.00 thinking |
| GPT-5.4 nano | OpenAI | $0.20 | $0.02 | $1.25 |
| DeepSeek V4-Flash (peak) | DeepSeek | $0.44 | $0.014 | $1.32 |
| Gemini 3.1 Flash-Lite | $0.25 | $0.025 | $1.50 | |
| Mistral Large 3 | Mistral AI | $0.50 | not published | $1.50 |
| DeepSeek V4-Pro (off-peak) | DeepSeek | $0.66 | $0.022 | $1.98 |
| GLM-4.7 | Z.ai | $0.60 | $0.11 | $2.20 |
| Gemini 2.5 Flash | $0.30 | $0.03 | $2.50 | |
| Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 | |
| Grok 4.3 (under 200K) | xAI | $1.25 | $0.20 | $2.50 |
| Gemini 3 Flash Preview | $0.50 | $0.05 | $3.00 | |
| GLM-5 | Z.ai | $1.00 | $0.20 | $3.20 |
| Gemini 3.8 Flash (through Dec 31, 2026) | $0.75 | $0.075 | $3.75 | |
| Gemini 3.7 Flash (through Dec 31, 2026) | $0.75 | $0.075 | $3.75 | |
| Gemini 3.6 Flash (through Dec 31, 2026) | $0.75 | $0.075 | $3.75 | |
| DeepSeek V4-Pro (peak) | DeepSeek | $1.32 | $0.044 | $3.96 |
| GLM-5.2 | Z.ai | $1.40 | $0.26 | $4.40 |
| GPT-5.4 mini | OpenAI | $0.75 | $0.075 | $4.50 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $0.10 | $5.00 |
| Qwen3-Max (0-32K) | Alibaba | $1.20 | not published | $6.00 |
| Grok 4.6 (under 200K) | xAI | $2.00 | $0.50 | $6.00 |
| Grok 4.5 (under 200K) | xAI | $2.00 | $0.30 | $6.00 |
| Mistral Medium 3.5 | Mistral AI | $1.50 | not published | $7.50 |
| Qwen3.7-Max | Alibaba | $2.50 | not published | $7.50 |
| Gemini 3.5 Flash | $1.50 | $0.15 | $9.00 | |
| Gemini 2.5 Pro (up to 200K) | $1.25 | $0.125 | $10.00 | |
| Claude Sonnet 5 | Anthropic | $2.00 | $0.20 | $10.00 |
| GPT-5.6 Terra | OpenAI | $2.00 | $0.20 | $12.00 |
| Gemini 3.1 Pro (up to 200K) | $2.00 | $0.20 | $12.00 | |
| GPT-5.4 | OpenAI | $2.50 | $0.25 | $15.00 |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $0.30 | $15.00 |
| Kimi K3 | Moonshot AI | $3.00 | $0.30 | $15.00 |
| GPT-5.6 Sol (promotional, at least through Nov 21, 2026) | OpenAI | $4.00 | $0.40 | $20.00 |
| Claude Opus 5 | Anthropic | $5.00 | $0.50 | $25.00 |
| GPT-5.5 | OpenAI | $5.00 | $0.50 | $30.00 |
| Claude Fable 5.1 | Anthropic | $10.00 | $0.25 | $50.00 |
| Claude Fable 5 | Anthropic | $10.00 | $1.00 | $50.00 |
| GPT-6 Astra | OpenAI | $10.00 | $1.00 | $50.00 |
| GPT-5.5 Pro | OpenAI | $30.00 | not published | $180.00 |
Alibaba publishes a third rate for its reasoning models that most comparisons omit. Qwen-Plus bills chain-of-thought output separately from the answer, at $4.00 per 1M tokens in the 0-256K band and $12.00 in the 256K-1M band, so a reasoning-mode request costs materially more than the non-thinking output rate suggests. Qwen3.7-Max also carried a limited-time 50 percent discount when these rates were checked on August 1, 2026.
One frontier vendor is missing from the table because it does not fit the table's rule. Meta's Muse Spark 1.3 is priced, according to The New Stack and VentureBeat's launch-day comparisons, at $1.25 input and $4.25 output on its standard tier, with a "Contributor" tier at $0.10 and $0.20 whose prompts Meta says it may use to improve its products.[34][35] Those figures are press-reported rather than read from a Meta price page and are included here for context only.
Sources by provider: OpenAI [1], Anthropic [6], Google [7], DeepSeek [8], xAI [9][39], Mistral [10], Alibaba [11], Z.ai [12], Moonshot [13].
Provider-by-provider detail
OpenAI
OpenAI's current families are GPT-6 Astra (released September 3, 2026), GPT-5.6 (Luna, Terra, Sol), the GPT-5.5 line, and the GPT-5.4 line. Cache reads are 90 percent off the base input rate. GPT-5.6 and GPT-6 Astra also charge for cache writes at 1.25x the uncached input rate, which arrived with GPT-5.6 and matches long-standing Anthropic practice.[1][3][4][33] As of September 3 the pricing page's headline table lists only Astra and the three GPT-5.6 models; the older lines are still priced on their individual model pages, and the GPT-5 family's model pages now mark their snapshots as deprecated.[1]
| Model | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
| gpt-6-astra | $10.00 | $1.00 | $12.50 | $50.00 |
| gpt-5.6-sol (promotional, at least through Nov 21, 2026) | $4.00 | $0.40 | $5.00 | $20.00 |
| gpt-5.6-terra | $2.00 | $0.20 | $2.50 | $12.00 |
| gpt-5.6-luna | $0.20 | $0.02 | $0.25 | $1.20 |
| gpt-5.6-cyber (Daybreak Red) | $12.50 | $1.25 | $15.625 | $75.00 |
| gpt-5.5 | $5.00 | $0.50 | not listed | $30.00 |
| gpt-5.5-pro | $30.00 | not published | not listed | $180.00 |
| gpt-5.4 | $2.50 | $0.25 | not listed | $15.00 |
| gpt-5.4-mini | $0.75 | $0.075 | not listed | $4.50 |
| gpt-5.4-nano | $0.20 | $0.02 | not listed | $1.25 |
| gpt-5.4-pro | $30.00 | not published | not listed | $180.00 |
| gpt-5.3-codex | $1.75 | $0.175 | not listed | $14.00 |
| gpt-5.2 | $1.75 | $0.175 | not listed | $14.00 |
| gpt-5.1 | $1.25 | $0.125 | not listed | $10.00 |
| gpt-5 (deprecated) | $1.25 | $0.125 | not listed | $10.00 |
| gpt-5-mini (deprecated) | $0.25 | $0.025 | not listed | $2.00 |
| gpt-5-nano (deprecated) | $0.05 | $0.005 | not listed | $0.40 |
| gpt-5-pro (deprecated) | $15.00 | not published | not listed | $120.00 |
The gpt-5.5-cyber row that appeared in the August 1 revision has been replaced: its model page no longer resolves, and the pricing page now lists gpt-5.6-cyber, the model behind the gpt-daybreak-red-latest alias, at the same $12.50 and $75 rates. The gpt-daybreak-blue-latest alias points to GPT-5.6 Sol at Sol's normal price.[1][2] The chat-latest alias, which tracks the model served in ChatGPT, is still priced at $5.00 and $30.00, so it is now more expensive than calling Sol directly.[1]
Three tier modifiers apply on top. Batch and flex both bill at half the standard rate; fast mode (renamed from priority processing on July 30, 2026) bills at double on GPT-6 Astra and on the GPT-5.6 and GPT-5.4 lines, and at 2.5 times standard on gpt-5.5. Regional processing endpoints for data residency add a 10 percent uplift on eligible models released on or after March 5, 2026, which includes Astra.[1][5][33]
| Model | Standard in/out | Batch and flex in/out | Fast mode in/out |
|---|---|---|---|
| gpt-6-astra | $10.00 / $50.00 | $5.00 / $25.00 | $20.00 / $100.00 |
| gpt-5.6-sol | $4.00 / $20.00 | $2.00 / $10.00 | $8.00 / $40.00 |
| gpt-5.6-terra | $2.00 / $12.00 | $1.00 / $6.00 | $4.00 / $24.00 |
| gpt-5.6-luna | $0.20 / $1.20 | $0.10 / $0.60 | $0.40 / $2.40 |
| gpt-5.5 | $5.00 / $30.00 | $2.50 / $15.00 | $12.50 / $75.00 |
| gpt-5.4 | $2.50 / $15.00 | $1.25 / $7.50 | $5.00 / $30.00 |
Anthropic
Anthropic publishes four prices per model rather than three, because a cache write is billed separately from a cache read and the write price depends on the cache lifetime. A 5-minute write costs 1.25x the base input rate, a 1-hour write costs 2x, and a read costs 0.1x, except on Claude Fable 5.1 and Claude Mythos 5.1, where a read costs 0.025x.[6]
| Model | Input | 5m cache write | 1h cache write | Cache hit | Output |
|---|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $12.50 | $20 | $0.25 | $50 |
| Claude Mythos 5.1 (limited availability) | $10 | $12.50 | $20 | $0.25 | $50 |
| Claude Fable 5 | $10 | $12.50 | $20 | $1.00 | $50 |
| Claude Mythos 5 (limited availability) | $10 | $12.50 | $20 | $1.00 | $50 |
| Claude Opus 5 | $5 | $6.25 | $10 | $0.50 | $25 |
| Claude Opus 4.8 | $5 | $6.25 | $10 | $0.50 | $25 |
| Claude Opus 4.7 | $5 | $6.25 | $10 | $0.50 | $25 |
| Claude Opus 4.6 | $5 | $6.25 | $10 | $0.50 | $25 |
| Claude Opus 4.5 | $5 | $6.25 | $10 | $0.50 | $25 |
| Claude Sonnet 5 | $2 | $2.50 | $4 | $0.20 | $10 |
| Claude Sonnet 4.6 | $3 | $3.75 | $6 | $0.30 | $15 |
| Claude Sonnet 4.5 | $3 | $3.75 | $6 | $0.30 | $15 |
| Claude Haiku 4.5 | $1 | $1.25 | $2 | $0.10 | $5 |
Two things changed between the two checks. The August 1 revision of this page listed Claude Sonnet 5 at $2 and $10 "through Aug 31, 2026" and at $3 and $15 "from Sep 1, 2026", which was what Anthropic had announced at launch. That increase did not happen: the pricing documentation now states that the $2 / $10 rate "is now the standard price" and that "the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur."[6] And Claude Fable 5.1 and Claude Mythos 5.1 appeared on the price list at the same $10 and $50 base rates as Fable 5 but with cache reads "priced at 0.025x the base input price", $0.25 per million tokens, against the 0.1x multiplier every other Claude model uses.[6][40] Fable 5.1's cache-write rates are unchanged at $12.50 and $20, so the discount applies only once a prefix is already cached.
Anthropic is the one major provider with no long-context surcharge: Claude 4.6 and later include the full 1,000,000-token window at standard rates, and the docs state that "a 900k-token request is billed at the same per-token rate as a 9k-token request."[6] The Batch API halves both input and output, which puts Fable 5.1 at $5 and $25 in batch. A fast mode research preview prices Claude Opus 5 and Opus 4.8 at $10 and $50, double standard, for "up to 2.5x faster speeds", and is not combinable with batch.[6][40] Requesting US-only inference with the inference_geo parameter on Claude 4.6 and later multiplies every token category by 1.1.[6]
Two line items sit outside the token price. Server-side web search is $10 per 1,000 searches on top of token costs. Claude Managed Agents adds $0.08 per session-hour of running time, metered only while a session is actually running.[6]
Google publishes four service tiers per model. Batch and flex are half the standard rate; priority is 1.8x. Context caching carries both a per-token read rate and an hourly storage charge, which is unusual and easy to miss when a cache sits idle.[7]
Google's Flash line has been repriced since August 1. Gemini 3.7 Flash, released August 13, launched at "an introductory price of half the original 3.6 Flash cost per million tokens", $0.75 input and $3.75 output, and the pricing page now applies that rate to Gemini 3.6 Flash as well; Gemini 3.8 Flash, released September 2 as Google's "third Flash release in only six weeks", launched "at the same introductory price as 3.7 Flash". All three revert to $1.50 and $7.50 on January 1, 2027, and their cache storage charge is $0.50 per hour during the promotion against $1.00 afterward.[7][37][41] The 3.6 Flash row in the August 1 revision ($1.50 / $7.50) was correct on that date and is superseded.
| Model | Input | Output | Cache read | Cache storage |
|---|---|---|---|---|
| Gemini 3.8 Flash (through Dec 31, 2026) | $0.75 | $3.75 | $0.075 | $0.50/hr |
| Gemini 3.7 Flash (through Dec 31, 2026) | $0.75 | $3.75 | $0.075 | $0.50/hr |
| Gemini 3.6 Flash (through Dec 31, 2026) | $0.75 | $3.75 | $0.075 | $0.50/hr |
| Gemini 3.8, 3.7, and 3.6 Flash (from Jan 1, 2027) | $1.50 | $7.50 | $0.15 | $1.00/hr |
| Gemini 3.5 Flash | $1.50 | $9.00 | $0.15 | $1.00/hr |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.03 | $1.00/hr |
| Gemini 3.1 Pro Preview (up to 200K) | $2.00 | $12.00 | $0.20 | $4.50/hr |
| Gemini 3.1 Pro Preview (over 200K) | $4.00 | $18.00 | $0.40 | $4.50/hr |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | $0.025 | $1.00/hr |
| Gemini 3 Flash Preview | $0.50 | $3.00 | $0.05 | $1.00/hr |
| Gemini 2.5 Pro (up to 200K) | $1.25 | $10.00 | $0.125 | $4.50/hr |
| Gemini 2.5 Pro (over 200K) | $2.50 | $15.00 | $0.25 | $4.50/hr |
| Gemini 2.5 Flash | $0.30 | $2.50 | $0.03 | $1.00/hr |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | $0.01 | $1.00/hr |
Audio input is priced above text, image, and video input on the Flash and Flash-Lite models: $1.00 against $0.30 on Gemini 2.5 Flash, $0.50 against $0.25 on Gemini 3.1 Flash-Lite, and $0.30 against $0.10 on Gemini 2.5 Flash-Lite.[7] As of September 3, 2026 the pricing page still lists no Pro-tier model above Gemini 3.1 Pro Preview.
DeepSeek
DeepSeek serves three model ids and, since August 16, 2026, publishes six prices for each: cache-hit input, cache-miss input, and output, each at a peak and an off-peak rate.[8]
| Model | Cache-hit input (off-peak / peak) | Cache-miss input (off-peak / peak) | Output (off-peak / peak) |
|---|---|---|---|
deepseek-v4-flash | $0.007 / $0.014 | $0.22 / $0.44 | $0.66 / $1.32 |
deepseek-v4-pro | $0.022 / $0.044 | $0.66 / $1.32 | $1.98 / $3.96 |
deepseek-v4-flash-vision-exp | $0.007 / $0.014 | $0.22 / $0.44 | $0.66 / $1.32 |
The rates in the August 1 revision of this page ($0.14 / $0.0028 / $0.28 for V4-Flash and $0.435 / $0.003625 / $0.87 for V4-Pro) were the prices DeepSeek charged at the time. On August 13, alongside the general-availability release of V4-Pro, DeepSeek's changelog announced that "with the official release of the DeepSeek V4 model family, we will update and adjust API pricing", adopting "peak/off-peak pricing, with off-peak prices set at half of the peak-hour prices", effective 16:00 UTC on August 16, 2026.[36] The off-peak V4-Flash rate is therefore about 57 percent higher on input and 2.4 times higher on output than the July rate, and the peak rate is double that again. The deepseek-v4-flash id now serves DeepSeek-V4-Flash-0731, deepseek-v4-pro serves DeepSeek-V4-Pro-0813, and the vision variant added on August 21 is billed at Flash rates with images converted to input tokens by their dimensions.[8][36]
The cache-hit rate is still the outlier of the whole market. At $0.007 against $0.22 off-peak, a hit costs about 3 percent of a miss, where the industry norm is 10 percent. Artificial Analysis called this out on July 31, under the earlier rates, describing a "~98% cache hit discount on its first-party API" against the 90 percent standard elsewhere.[25]
xAI
xAI prices every model in two bands split at 200,000 tokens, with the upper band exactly double the lower on all three rates.[9] Since these rows were checked on August 1, xAI has released Grok 4.6, and as of September 3 its models page presents Grok 4.6 alone as "our flagship model for code and everything else" at $2.00 input, $0.50 cached input, and $6.00 output on a 500,000-token window, with the model page stating that requests above 200,000 tokens are charged at different rates and that the Batch API is not supported.[9][39] VentureBeat's launch-day table gives the upper band as $4.00 and $12.00.[35] The older Grok rows below are as checked on August 1 and were not re-verified.
| Model | Context | Input (under 200K) | Cached (under 200K) | Output (under 200K) | Input (200K+) | Output (200K+) |
|---|---|---|---|---|---|---|
| Grok 4.6 | 500K | $2.00 | $0.50 | $6.00 | $4.00 (press-reported) | $12.00 (press-reported) |
| Grok 4.5 | 500K | $2.00 | $0.30 | $6.00 | $4.00 | $12.00 |
| Grok 4.3 | 1M | $1.25 | $0.20 | $2.50 | $2.50 | $5.00 |
| Grok 4.20-0309 (reasoning) | 1M | $1.25 | $0.20 | $2.50 | $2.50 | $5.00 |
| Grok 4.20-0309 (non-reasoning) | 1M | $1.25 | $0.20 | $2.50 | $2.50 | $5.00 |
| Grok 4.20 multi-agent 0309 | 1M | $1.25 | $0.20 | $2.50 | $2.50 | $5.00 |
| Grok build 0.1 | 256K | $1.00 | $0.20 | $2.00 | $2.00 | $4.00 |
Grok 4.6's cached-input rate is a smaller discount than its predecessor's: $0.50 against $2.00 is 75 percent off, where Grok 4.5's $0.30 was 85 percent off.[9][39]
Mistral AI
Mistral does not publish per-model cache-hit rates, stating a flat 90 percent saving on cached input tokens, and offers 50 percent off for batch processing.[10]
| Model | Input | Output |
|---|---|---|
| Mistral Medium 3.5 | $1.50 | $7.50 |
| Mistral Large 3 | $0.50 | $1.50 |
| Mistral Small 4 | $0.15 | $0.60 |
| Ministral 3 (3B, 8B, 14B) | $0.10 to $0.20 | $0.10 to $0.20 |
| Magistral Medium | $2.00 | $5.00 |
| Magistral Small | $0.50 | $1.50 |
| Codestral | $0.30 | $0.90 |
| Devstral 2 | $0.40 | $2.00 |
The Ministral 3 line is priced symmetrically, the same rate for input and output, which is rare and makes it unusually attractive for output-heavy chat where responses run long.
Alibaba (Qwen)
Alibaba's international Model Studio endpoint prices most Qwen models in input-length bands, which is a different mechanism from OpenAI's or Google's single threshold: the band is chosen by the request's input length and applies to that request.[11]
| Model | Input band | Input | Output |
|---|---|---|---|
| Qwen3.7-Max | 0 to 1M | $2.50 | $7.50 |
| Qwen3-Max | 0 to 32K | $1.20 | $6.00 |
| Qwen3-Max | 32K to 128K | $2.40 | $12.00 |
| Qwen3-Max | 128K to 256K | $3.00 | $15.00 |
| Qwen-Plus | 0 to 256K | $0.40 | $1.20 |
| Qwen-Plus | 256K to 1M | $1.20 | $3.60 |
| Qwen-Flash | 0 to 256K | $0.05 | $0.40 |
| Qwen-Flash | 256K to 1M | $0.25 | $2.00 |
Qwen-Flash at $0.05 input ties the now-deprecated GPT-5-nano for the lowest published input rate of any model in this article. Note that earlier revisions of this page listed Qwen3-Max at $0.86 and $3.44; those figures do not match Alibaba's current international price list and have been corrected.
Z.ai (GLM)
| Model | Input | Cached input | Output |
|---|---|---|---|
| GLM-5.2 | $1.40 | $0.26 | $4.40 |
| GLM-5 | $1.00 | $0.20 | $3.20 |
| GLM-4.7 | $0.60 | $0.11 | $2.20 |
| GLM-4.5-Air | $0.20 | $0.03 | $1.10 |
Z.ai lists cached-input storage fees as free for a limited time on all four models.[12]
Moonshot AI (Kimi)
Moonshot prices Kimi K3 flat across its full 1,048,576-token window, with the discounted input rate applied automatically when prefix caching hits.[13]
| Model | Cache-hit input | Uncached input | Output |
|---|---|---|---|
| Kimi K3 | $0.30 | $3.00 | $15.00 |
The July 30 price cut: Luna down 80 percent, Terra down 20 percent
The most consequential pricing event of late July was OpenAI's July 30 repricing. The API changelog records it plainly: "Starting July 30, GPT-5.6 Luna costs 80% less, while GPT-5.6 Terra costs 20% less."[2] The pricing page then listed Luna at $0.20 input, $0.02 cached input, and $1.20 output, against the $1.00 / $0.10 / $6.00 it carried at launch three weeks earlier. Terra fell from $2.50 / $0.25 / $15.00 to $2.00 / $0.20 / $12.00. GPT-5.6 Sol was not cut that day and stayed at $5.00 / $0.50 / $30.00 until its own cut three weeks later, covered in the next section.[1][2][17]
The shape of the cut is as interesting as its size. OpenAI repriced the two cheaper tiers and left the flagship untouched, which widened the spread inside a single model family from 5x to 25x on both input and output. Nothing else changed: the long-context surcharge above 272,000 input tokens is still 2x input and 1.5x output, cache writes are still 1.25x the uncached input rate, and rate limits were not adjusted.[1][3]
The effect on the rankings is real. Luna is now cheaper on output than GPT-5.4 nano, a model two tiers below it in capability, and cheaper on both rates than Gemini 2.5 Flash. At one twenty-fifth of Sol's then-headline price, a model that Artificial Analysis scored at 51 on its Intelligence Index now sits in the same price band as the industry's smallest models.[18][20]
A second pricing change shipped the same day, covered under service tiers below.
The August 21 Sol cut and the GPT-6 Astra launch price
OpenAI's next two pricing moves came in opposite directions. On August 21 the changelog recorded that "GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens, representing 20% lower input pricing and 33% lower output pricing", with cached input at $0.40 and cache writes at $5.00, and added that "GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026."[1][2] Every tier moved with it: batch and flex to $2 / $10, fast mode to $8 / $40, and the long-context band above 272,000 input tokens to $8 / $30.[1][30] The word "promotional" matters. OpenAI has not said what Sol will cost after November 21, and the earlier $5 / $30 rate that this page carried until this revision should be read as the pre-promotion list price, not as a current one. The cut narrowed the spread inside the GPT-5.6 family from 25x to 20x on input and about 17x on output.
Thirteen days later, on September 3, OpenAI announced GPT-6 Astra and published its API price as gpt-6-astra: $10 input, $1 cached input, $12.50 cache write, and $50 output per million tokens.[1][32][33] The launch post gives the two headline numbers and says only that "separate rates apply to cache reads and writes"; the pricing page fills in the rest, and notes that the initial rollout is to enterprises in OpenAI's Trusted Access Program with API access "coming in the coming days".[1][32] The full ladder is:
| GPT-6 Astra tier | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
| Standard (up to 272K input) | $10.00 | $1.00 | $12.50 | $50.00 |
| Standard, long context (over 272K input) | $20.00 | $2.00 | $25.00 | $75.00 |
| Batch and Flex | $5.00 | $0.50 | $6.25 | $25.00 |
| Fast mode | $20.00 | $2.00 | $25.00 | $100.00 |
| Fast mode, long context | $40.00 | $4.00 | $50.00 | $150.00 |
The mechanics are the GPT-5.6 mechanics at a higher base. The model page lists a 1,050,000-token context window and 128,000 maximum output tokens, and prompts above 272,000 input tokens are "priced at 2x input and cache rates and 1.5x output for the full request", cache writes are 1.25x uncached input, batch and flex are half, and fast mode is double.[33] Fast mode, in the launch post as first published, "delivers up to 2.5x the speed of Standard processing at 2x the Standard price" (a copy of the post fetched on September 4 says "up to 2x the speed"), carries no latency SLA for Astra (unlike GPT-5.6 and earlier), and is unavailable for Astra with EU data residency; regional endpoints otherwise add the usual 10 percent.[1][30][32] Two API restrictions affect spend rather than price. Astra "does not support the none reasoning effort level", so unlike the GPT-5.6 models it cannot be run without reasoning tokens, and its supported effort settings run from low to max.[2][33] OpenAI's launch post as first published said Astra "is also available in Amazon Bedrock" (a copy fetched on September 4 says it "will be available" there); the pricing page notes that "OpenAI models in Amazon Bedrock are billed through AWS and may differ from direct OpenAI pricing", and when checked on September 3 the Amazon Bedrock pricing page listed OpenAI's GPT-5.6, GPT-5.5, GPT-5.4, and Daybreak models but no Astra row yet.[1][16][32]
Against the rest of the market, Astra's list price is 2.5 times Sol's promotional rate on both input and output, and, as The New Stack noted, "it matches Anthropic's pricing for Fable 5.1."[34] That match is exact on the base rates and on the 5-minute cache write ($12.50 at both), but not on cache reads: a cached token costs $1.00 on Astra and $0.25 on Fable 5.1, so a heavily cached agent loop is cheaper on Anthropic's side of the tie.[1][6] TNS also set the price against the cheaper frontier tier, "far above" Muse's $1.25 / $4.25 and Gemini 3.8 Flash's introductory $0.75 / $3.75.[34]
OpenAI's answer to that comparison is the argument covered in the next section. "The price per task is what matters," Greg Brockman told reporters at the launch briefing, and VentureBeat headed its pricing section "Price-per-task now matters more than price-per-token, according to OpenAI."[34][35] The company's supporting numbers are its own: in latency simulations on OSWorld 2.0, it reports Astra scoring 72.6 percent at roughly 40 minutes per task against Sol's 65.7 percent at roughly 75 minutes, "about 47% less time per task", and it says Astra used "substantially fewer output tokens" than Sol on its ExploitGym evaluation.[32] Those are company-reported figures from the launch post, not independent measurements, and TNS's assessment was that "the launch data is too sparse to show whether those savings offset the price premium."[34]
Why headline per-token prices mislead
This is the part naive comparisons get wrong, and it now dominates the arithmetic.
Reasoning models spend output tokens thinking before they answer, and how many they spend is a setting, not a constant. Willison, who had early access to GPT-5.6, priced one identical SVG-drawing prompt across all three models at six effort levels. The cheapest run was Luna at effort none for 0.71 cents. The most expensive was Sol at max for 48.55 cents. Same prompt, same family, same week, a spread of about 68 times, most of it from token counts rather than rates. His conclusion: "price-per-million tokens doesn't tell us much now that the number of reasoning tokens can differ so much between models for the same task."[17] Those figures predate the July 30 cut and the August 21 Sol cut, so both ends of that range are now lower.
The better comparison is cost per task. Artificial Analysis measures the full dollar cost of running its Intelligence Index on each model, computing each evaluation's cost "from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight", and using the token counts each provider's own API reports rather than a standardized client-side tokenizer.[24][27] The figures below are from Artificial Analysis model pages as listed on August 1, 2026, all at maximum reasoning effort, with GPT-6 Astra added from its page as listed on September 3.
| Model | Intelligence Index | Cost to run the full index | Blended rate (7:2:1) |
|---|---|---|---|
| GPT-6 Astra (September 3) | 61 | $3,013.30 | not listed |
| Claude Fable 5 | 60 | $5,630.52 | $7.70 |
| GPT-5.6 Sol (at the pre-August 21 rate) | 59 | $3,442.81 | $4.35 |
| GPT-5.6 Luna | 51 | $190.87 | $0.17 |
| DeepSeek V4-Flash 0731 (at the pre-August 16 rate) | 50 | $72.02 | $0.06 |
| DeepSeek V4-Pro (at the pre-August 16 rate) | 44 | $176.34 | $0.18 |
Per-model sources: GPT-6 Astra [38], Claude Fable 5 [23], GPT-5.6 Sol [19], GPT-5.6 Luna [20], DeepSeek V4-Flash [21], DeepSeek V4-Pro [22]. The blended rate is Artificial Analysis's own 7:2:1 weighting of cache-hit, input, and output tokens. The Sol and DeepSeek rows were computed under prices that have since changed and are kept as measured; they have not been rescaled.
Astra's row is the first independent cost figure for the new model. On the day of launch Artificial Analysis listed GPT-6 Astra at maximum effort with an Intelligence Index score of 61, a cost of $3,013.30 to run the full index, $1.67 per index task, and 42 million output tokens generated over the evaluation, which it described as "fairly concise in comparison to the median of 62M", while summarizing the model as "amongst the leading models in intelligence, but particularly expensive when comparing to other models of similar price."[38] Read against the table, that puts Astra's measured cost below Claude Fable 5's at a higher index score, and below the figure Sol posted in August at its old $5 / $30 list price, which is the shape of the price-per-task argument OpenAI is making; it also says nothing yet about the workloads a buyer actually runs.
Three things fall out of that table that a per-token comparison hides.
Cost per task compresses the headline gap. Sol's list rates were 25 times Luna's on both input and output when these runs were made, but the measured cost of running the same benchmark suite differed by about 18 times. Luna talks more: Artificial Analysis noted it generated 130 million output tokens over the evaluation.[19][20] Verbosity ate roughly a quarter of the apparent saving.
Cost per task can invert the headline ranking. DeepSeek V4-Pro carries a higher list price than V4-Flash and scored six points lower on the index, yet cost more than twice as much to evaluate, because its per-token rates are roughly three times higher. It was not the more verbose of the two: Artificial Analysis recorded 180 million generated tokens for V4-Pro against 210 million for V4-Flash-0731.[21][22]
Similar scores can sit at very different prices. Artificial Analysis put V4-Flash 0731's cost per task at "~60% lower than GPT-5.6 Luna (max), a model with comparable intelligence", and that was measured the day after OpenAI's 80 percent cut.[25] The full-index figures agree: $72.02 against $190.87, a 62 percent difference, for scores of 50 and 51. Within days that spread was general business-press news: Bloomberg reported on August 4, 2026, citing Artificial Analysis benchmark testing, that executing a complex real-world workload cost $0.03 with DeepSeek V4-Flash against $3.15 with Claude Fable 5, and wrote that a so-called DeepSeek death zone had emerged on a widely shared Artificial Analysis chart for rivals that charge more for the same capability or deliver less for the same price.[31] DeepSeek's August 16 repricing, which raised V4-Flash's off-peak output rate 2.4x, postdates all of those measurements.
For a sense of the top of the range, Artificial Analysis reported in June that Claude Fable 5 was "the most expensive model we have ever benchmarked", at roughly 1.7 times the next-highest model, Claude Opus 4.8, and 2.2 times GPT-5.5 at extra-high effort.[26] The July article on the GPT-5.6 launch gave per-task costs of $1.04 for Sol at max effort, $0.55 for Terra, and $0.21 for Luna, and observed that Sol offered "a similar level of intelligence to Claude Fable 5 at approximately one third of the cost."[18][29] Those per-task figures also predate the Luna and Sol cuts.
The practical upshot: benchmark your own prompt at the effort level you will actually ship, and compare total spend. A rate card cannot tell you how talkative a model is.
Cache pricing is now a first-order cost
Context caching has moved from a nice-to-have to one of the largest levers on a real bill, especially for agents that resend a growing conversation on every turn. Three separate prices are involved, and providers differ on all three.
| Provider | Cache-hit discount | Cache-write charge | Storage charge |
|---|---|---|---|
| DeepSeek | ~97 percent ($0.007 vs $0.22 off-peak on V4-Flash) | none published | none published |
| OpenAI | 90 percent | 1.25x uncached input on GPT-5.6 and GPT-6 Astra | none published |
| Anthropic | 90 percent (0.1x base input); 97.5 percent (0.025x) on Claude Fable 5.1 and Mythos 5.1 | 1.25x for 5 minutes, 2x for 1 hour | none published |
| 90 percent | none published | $0.50 to $4.50 per hour depending on model | |
| xAI | 75 percent on Grok 4.6, 85 percent on Grok 4.5, 84 percent on Grok 4.3 | none published | none published |
| Mistral | 90 percent (stated, not per model) | none published | none published |
| Z.ai | 80 to 85 percent depending on model | none published | free for a limited time |
| Moonshot | 90 percent | none published | none published |
Sources: [1][6][7][8][9][10][12][13][39].
Two consequences. First, DeepSeek's cache economics are structurally different from everyone else's, not marginally better: at about 97 percent off, a 100,000-token prefix resent a thousand times costs $0.70 in cache hits instead of $22.00 in cache misses at off-peak rates. Artificial Analysis attributed much of DeepSeek's cost-per-task advantage to exactly this.[25] Anthropic has now moved its top model most of the way toward that structure: on Claude Fable 5.1 the same 100 million cached tokens cost $25 against $1,000 uncached, where on GPT-6 Astra, at the same base price, they cost $100.[1][6] Second, a cache is not free to fill. Anthropic's own guidance is that a 5-minute cache pays for itself after one read and a 1-hour cache after two, which means a cache with a low hit rate can cost more than no cache at all.[6] Google's hourly storage charge creates the same trap in a different shape: an idle cached context on Gemini 3.1 Pro accrues $4.50 an hour whether or not anything reads it.[7] See prompt caching for the mechanics.
Long-context surcharges
Several providers reprice an entire request once its prompt crosses a threshold. This is not a marginal rate on the tokens above the line; at OpenAI it applies to the full request.
| Provider | Threshold | What changes |
|---|---|---|
| OpenAI (GPT-6 Astra, GPT-5.6, 5.5, 5.4) | 272,000 input tokens | 2x input and cache rates and 1.5x output for the full request[3][4][33] |
| Google (Gemini 3.1 Pro, 2.5 Pro) | 200,000 tokens | Input and cache roughly double, output rises 1.5x on 3.1 Pro and 1.5x on 2.5 Pro[7] |
| xAI (all models) | 200,000 tokens | All three rates double[9] |
| Alibaba (Qwen) | 32K, 128K, 256K bands by model | Rates step up per band[11] |
| Anthropic (Claude 4.6+) | none | Full 1M window at standard rates[6] |
| DeepSeek, Moonshot | none published | Flat to 1M tokens[8][13] |
The practical reading: Anthropic and DeepSeek are the cheapest places to put a genuinely enormous prompt relative to their own short-context rates, while a GPT-5.6 Sol request at 300,000 input tokens is billed at $8 input and $30 output, not $4 and $20, and the same request on GPT-6 Astra is billed at $20 and $75, not $10 and $50. For window sizes rather than prices, see LLM context window comparison.
Batch, flex, and premium tiers
Nearly every provider sells the same model at three or four latency classes.
| Tier | Typical price | Trade-off |
|---|---|---|
| Batch | 50 percent of standard at OpenAI, Anthropic, Google, and Mistral | Asynchronous, results in hours[1][6][7][10] |
| Flex | 50 percent of standard at OpenAI and Google | Synchronous but slower, with occasional 429 resource-unavailable responses[5][7] |
| Standard | list | baseline |
| Fast or priority | 2x at OpenAI and Anthropic, 1.8x at Google | Lower latency[1][6][7] |
OpenAI now has the cleanest version of this ladder in the market, and the second half of its July 30 changes built it. Priority processing was renamed Fast mode that day, a rename recorded in a footnote on the pricing page and repeated in the guide: "Priority processing was renamed Fast mode on July 30, 2026."[1][30] Fast mode is exactly double standard on every GPT-5.6 tier and on GPT-6 Astra, giving a four-step ladder at 0.5x, 0.5x, 1x, and 2x of list:
| Model | Batch | Flex | Standard | Fast mode |
|---|---|---|---|---|
| gpt-6-astra | $5.00 / $25.00 | $5.00 / $25.00 | $10.00 / $50.00 | $20.00 / $100.00 |
| gpt-5.6-sol | $2.00 / $10.00 | $2.00 / $10.00 | $4.00 / $20.00 | $8.00 / $40.00 |
| gpt-5.6-terra | $1.00 / $6.00 | $1.00 / $6.00 | $2.00 / $12.00 | $4.00 / $24.00 |
| gpt-5.6-luna | $0.10 / $0.60 | $0.10 / $0.60 | $0.20 / $1.20 | $0.40 / $2.40 |
The documented speed gain is "up to 2.5x faster than Standard processing" for gpt-5.6-sol and "up to 2.5x the speed of Standard processing" for gpt-6-astra; OpenAI publishes no figure for Terra or Luna. Since August 5, Fast mode also accepts long-context requests above 272,000 tokens on the GPT-5.6 models. The rename is backward compatible, and either service_tier: "priority" or service_tier: "fast" selects the tier.[2][30][32] A fifth rung is coming: on August 13 OpenAI announced "Ultrafast mode", a service tier for GPT-5.6 Sol "that runs up to 14x faster than Standard processing", in limited preview to select customers with no published price.[2]
OpenAI's flex tier bills "at Batch API rates, with additional discounts from prompt caching", is in beta with limited model availability, and needs a raised client timeout because request timeouts are more likely; requests that fail with 429 are not charged.[5] Anthropic's fast mode is a research preview limited to Claude Opus 5 and Opus 4.8, is first-party only, and cannot be combined with the Batch API.[6] For any non-interactive job, evaluations, bulk extraction, data labeling, the batch tier halves every number in this article. See the OpenAI Batch API page for mechanics.
Peak and off-peak pricing
DeepSeek is the one provider covered here that charges by the clock, and it has done so since 16:00 UTC on August 16, 2026. Its pricing page states that "off-peak rates are half of the peak rates" and defines peak hours as "01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak)", which is 09:00-12:00 and 14:00-18:00 Beijing time on weekdays.[8] Weekends are entirely off-peak.
The August 1 revision of this page reported the peak policy as announced but not in force, which was accurate then: DeepSeek's page said the effective date "will be subject to the official announcement", and no date had been published as of August 5. The announcement came in the August 13 changelog entry for V4-Pro's general availability, and the schedule has been live since August 16.[36] Any comparison built on DeepSeek's July rates, or on a single DeepSeek rate without saying which hours it applies to, is now wrong. For a job that can wait, the difference is a factor of two: a million output tokens on V4-Flash cost $0.66 at 12:00 UTC on a Monday and $1.32 at 09:00.
No other provider covered here publishes time-of-day pricing.
Open weights change what an API price means
For a model published under a permissive license, the API rate is a price for a service, not a price for the model. DeepSeek V4 is the clearest case: DeepSeek-V4-Pro carries an MIT license tag on Hugging Face, with 1.6 trillion total parameters, 49 billion active, and a one-million-token context window.[28] You can run it on your own hardware and pay nothing per token, or buy it from someone other than DeepSeek.
The spread between hosts for identical weights is large:
| Model | First-party API | Third-party host |
|---|---|---|
| DeepSeek V4-Pro | $0.66 / $1.98 off-peak, $1.32 / $3.96 peak (DeepSeek, September 3)[8] | $1.74 / $3.48, cached $0.20 (Together AI, August 1)[14] |
| Kimi K3 | $3.00 / $15.00 (Moonshot)[13] | $3.00 / $15.00, cached $0.30 (Together AI)[14] |
| GLM-5.2 | $1.40 / $4.40 (Z.ai)[12] | $1.40 / $4.40, cached $0.26 (Together AI)[14] |
When both were checked on August 1, Together AI charged exactly four times DeepSeek's own rate for V4-Pro while matching first-party pricing on Kimi K3 and GLM-5.2. DeepSeek's August 16 repricing closed most of that gap from DeepSeek's side: against Together's August 1 rate, which was not re-checked, the first-party API is now about 2.6 times cheaper off-peak and 1.3 times cheaper at peak on input, and slightly more expensive than Together on output during peak hours ($3.96 against $3.48). Speed-oriented hosts price differently again. Groq lists GPT-OSS 120B at $0.15 and $0.60, GPT-OSS 20B at $0.075 and $0.30, Llama 3.1 8B Instant at $0.05 and $0.08, Llama 3.3 70B Versatile at $0.59 and $0.79, and Qwen 3.6 27B at $0.60 and $3.00.[15] Amazon Bedrock prices its own Amazon Nova family at $0.035 and $0.14 for Nova Micro, $0.06 and $0.24 for Nova Lite, and $0.80 and $3.20 for Nova Pro in the US East and US West regions, and hosts third-party open models such as DeepSeek V3.2 at $0.62 and $1.85 and Qwen3 Next 80B at $0.15 and $1.20.[16] Nova Micro's $0.035 input rate is the lowest per-token input price of any model named in this article.
Aggregators including OpenRouter resell many of these endpoints, and their listed prices track the underlying host rather than the model. Availability, rate limits, and feature parity move independently of the model release, so a price quoted for "DeepSeek V4-Pro" is meaningless without naming who is serving it, and, since August 16, at what hour.
Tokenizers, and other reasons per-token rates are not comparable
A per-token price is only comparable if the token counts are. Anthropic states that Claude 4.7 and later models, plus Claude Mythos Preview, use a newer tokenizer that "produces approximately 30% more tokens for the same text", with the exact increase depending on content and workload; Claude Sonnet 4.6 and earlier use the previous tokenizer.[6] That erodes a meaningful part of any apparent saving from a lower headline rate on a newer Claude model, and it means comparing Opus 5 against Sonnet 4.6 on rate alone understates Opus 5's cost. It also means the GPT-6 Astra and Claude Fable 5.1 tie at $10 and $50 is a tie in rate only; the two models will not produce the same token count from the same text. Artificial Analysis handles the same problem by using each provider's reported token counts for cost reporting while standardizing on a single tokenizer for its speed measurements.[27] See tokenization for the underlying mechanics.
Two smaller distortions are worth naming. Output tokens run about 2x to 8x input across the models in this article, so output price dominates chat and generation workloads while input price dominates retrieval, classification, and long-document reading; that is why the main table is sorted by output. And several charges sit entirely outside the token meter, including Anthropic's $10 per 1,000 web searches and $0.08 per session-hour for Managed Agents, and Google's hourly cache storage.[6][7]
How these prices were verified
Every figure was read on August 1, 2026 from the provider's own pricing documentation: OpenAI [1], Anthropic [6], Google [7], DeepSeek [8], xAI [9], Mistral [10], Alibaba Cloud Model Studio [11], Z.ai [12], Moonshot [13], Together AI [14], Groq [15], and Amazon Web Services [16]. On September 3, 2026, the OpenAI, Anthropic, Google, and DeepSeek pages, every OpenAI model page listed above, the xAI models page and Grok 4.6 model page, and the Amazon Bedrock pricing page were read again. The Mistral, Alibaba, Z.ai, Moonshot, Together AI, and Groq rows, and the older xAI rows, were not re-checked and stand as of August 1. Cost-per-task and cost-to-run figures come from Artificial Analysis model pages and articles, dated where the date matters.
No number here was carried forward from the previous revision without being re-checked at least once, and both checks found real drift. The August 1 check found Qwen3-Max listed at $0.86 and $3.44 against Alibaba's current $1.20 and $6.00 base band; Grok 4.3 shown without its published $0.20 cache rate and without the 200,000-token doubling; Grok 4.1 Fast gone from xAI's model page; and the whole GPT-5.6 family, Claude Opus 5, Claude Sonnet 5, Gemini 3.6 Flash, Kimi K3, and the GLM line absent. The September 3 check found GPT-5.6 Sol on a promotional rate a third below the one this page carried; GPT-6 Astra, Claude Fable 5.1, Claude Mythos 5.1, Gemini 3.7 Flash, Gemini 3.8 Flash, and Grok 4.6 missing; Gemini 3.6 Flash halved; DeepSeek on new base rates with peak billing in force where this page had said it was not; Claude Sonnet 5's scheduled increase cancelled; the gpt-5.5-cyber model page gone and gpt-5.6-cyber in its place; and the GPT-5 family flagged as deprecated. Third-party aggregators were used only to locate official pages, never as a price source, and the two press-reported figures (Meta's Muse Spark rates and Grok 4.6's long-context band) are labelled as such. Where a current price could not be confirmed on an official page, the model was omitted.
Related pages: AI pricing for subscription and market economics, inference for what you are actually paying for, frontier models for the capability tiers referenced here, and LLM benchmark comparison for the capability side of the price-performance question.
References
- ^OpenAI, "Pricing," OpenAI API docs (accessed August 1 and September 3, 2026). developers.openai.com/...pricing
- ^OpenAI, "Changelog," OpenAI API docs (accessed August 1 and September 3, 2026). developers.openai.com/...changelog
- ^OpenAI, "gpt-5.6-luna," OpenAI API model reference (accessed August 1, 2026). developers.openai.com/...gpt-5.6-luna
- ^OpenAI, "gpt-5.6," OpenAI API model reference (accessed August 1 and September 3, 2026). developers.openai.com/...gpt-5.6
- ^OpenAI, "Flex processing," OpenAI API docs (accessed August 1, 2026). developers.openai.com/...flex-processing
- ^Anthropic, "Pricing," Claude Platform docs (accessed August 1 and September 3, 2026). platform.claude.com/...pricing
- ^Google, "Gemini Developer API pricing," Google AI for Developers (accessed August 1 and September 3, 2026). ai.google.dev/...pricing
- ^DeepSeek, "Models & Pricing," DeepSeek API docs (accessed August 1 and September 3, 2026). api-docs.deepseek.com/...pricing
- ^xAI, "Models," xAI docs (accessed August 1 and September 3, 2026). docs.x.ai/...models
- ^Mistral AI, "API pricing," La Plateforme (accessed August 1, 2026). mistral.ai/...api
- ^Alibaba Cloud, "Model pricing," Model Studio documentation (accessed August 1, 2026). alibabacloud.com/...model-pricing
- ^Z.ai, "Pricing," Z.ai developer guides (accessed August 1, 2026). docs.z.ai/...pricing
- ^Kimi, "Kimi K3 Pricing Explained: Plans and API Costs" (accessed August 1, 2026). kimi.com/...kimi-k3-pricing
- ^Together AI, "Pricing" (accessed August 1, 2026). together.ai/pricing
- ^Groq, "Pricing" (accessed August 1, 2026). groq.com/pricing
- ^Amazon Web Services, "Amazon Bedrock pricing" (accessed August 1 and September 3, 2026). aws.amazon.com/...pricing
- ^Simon Willison, "The new GPT-5.6 family: Luna, Terra, Sol," Simon Willison's Weblog, July 9, 2026. simonwillison.net/...gpt-5-6
- ^Artificial Analysis, "GPT-5.6 has landed," July 9, 2026. artificialanalysis.ai/...gpt-5-6-has-landed
- ^Artificial Analysis, "GPT-5.6 Sol: Intelligence, Performance & Price Analysis" (accessed August 1, 2026). artificialanalysis.ai/...gpt-5-6-sol
- ^Artificial Analysis, "GPT-5.6 Luna: Intelligence, Performance & Price Analysis" (accessed August 1, 2026). artificialanalysis.ai/...gpt-5-6-luna
- ^Artificial Analysis, "DeepSeek V4 Flash: Intelligence, Performance & Price Analysis" (accessed August 1, 2026). artificialanalysis.ai/...deepseek-v4-flash
- ^Artificial Analysis, "DeepSeek V4 Pro: Intelligence, Performance & Price Analysis" (accessed August 1, 2026). artificialanalysis.ai/...deepseek-v4-pro
- ^Artificial Analysis, "Claude Fable 5: Intelligence, Performance & Price Analysis" (accessed August 1, 2026). artificialanalysis.ai/...claude-fable-5
- ^Artificial Analysis, "Models" index (accessed August 1, 2026). artificialanalysis.ai/models
- ^Artificial Analysis, "DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index," @ArtificialAnlys on X, July 31, 2026. x.com/...2083123180869496865
- ^Artificial Analysis, "Claude Fable 5 cost ~$6.2K to run the Artificial Analysis Intelligence Index benchmarks," @ArtificialAnlys on X, June 17, 2026. x.com/...2067384319942029379
- ^Artificial Analysis, "Intelligence benchmarking methodology" (accessed August 1, 2026). artificialanalysis.ai/...intelligence-benchmarking
- ^DeepSeek, "DeepSeek-V4-Pro," Hugging Face model repository (accessed August 1, 2026). huggingface.co/...DeepSeek-V4-Pro
- ^Artificial Analysis, "How GPT-5.6 Sol, Terra, Luna compare on intelligence vs cost," July 13, 2026. artificialanalysis.ai/...ost-across-sol-terra-luna
- ^OpenAI, "Fast mode," OpenAI API docs (accessed August 1 and September 3, 2026). developers.openai.com/...fast-mode
- ^Zheping Huang, Nectar Gan and Saritha Rai, "China's AI Blitz Creates 'Death Zone' for Rival US Model Makers," Bloomberg News, August 4, 2026. bloomberg.com/...th-zone-for-rival-us-model-makers
- ^OpenAI, "GPT-6 Astra: A new generation of intelligence," September 3, 2026 (the page returned a 404 for part of the afternoon of publication; text checked against contemporaneous press reports and a reader-saved copy). openai.com/...gpt-6-astra
- ^OpenAI, "gpt-6-astra," OpenAI API model reference (accessed September 3, 2026). developers.openai.com/...gpt-6-astra
- ^Frederic Lardinois, "OpenAI launches GPT-6 Astra and says welcome to the 'AGI era'," The New Stack, September 3, 2026. thenewstack.io/openai-gpt6-astra-benchmarks
- ^Carl Franzen, "'Welcome to the AGI era': OpenAI launches GPT-6 Astra," VentureBeat, September 3, 2026. venturebeat.com/...era-openai-launches-gpt-6-astra
- ^DeepSeek, "Change Log," DeepSeek API docs, entries dated 2026-08-13 and 2026-08-21 (accessed September 3, 2026). api-docs.deepseek.com/updates
- ^Google, "Introducing Gemini 3.7 Flash," The Keyword, August 13, 2026. blog.google/...introducing-gemini-3-7-flash
- ^Artificial Analysis, "GPT-6 Astra: Intelligence, Performance & Price Analysis" (accessed September 3, 2026). artificialanalysis.ai/...gpt-6-astra
- ^xAI, "Grok 4.6," xAI docs model page (accessed September 3, 2026). docs.x.ai/...grok-4.6
- ^Anthropic, "Pricing," anthropic.com (accessed September 3, 2026). anthropic.com/pricing
- ^Google, "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber," The Keyword, September 2, 2026. blog.google/...3-8-flash-and-3-8-flash-cyber
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
7 revisions · v8 · 9,833 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independently fact-checked on September 3-4, 2026 against the cited primary sources (OpenAI, benchmark maintainers, vendor pricing pages) and press; verifier findings applied before publication.
Cite this page: AI Wiki. "LLM API Pricing Comparison." aiwiki.ai, updated 4 Sept 2026, fact-checked 4 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/llm_api_pricing_comparison