Split input correctly
Cached tokens are carved out of total input. In Standard mode, only that subset receives a recorded cache rate; Batch uses its own published input rate. Published cache-write fees are disclosed, not silently added.
Turn a token workload into per-request, monthly, and annual estimates. Compare official prices and see which model context windows can actually hold the request.
Private, browser-only estimate
No API key is needed. Your workload inputs stay on this device and are not saved or sent to a provider. The address bar carries your settings so the estimate can be shared.
Workload
Cached input is a subset of input, never an additional token count.
Presets (illustrative defaults, not provider guidance)
Per request
$0.0688
22K total tokens
Per month
$1,512.50
22K requests
Per year
$18,150.00
12 equal months
Cache writes bill at $3.125 per 1M and are not modeled here. OpenAI bills prompt-cache writes for this model at 1.25x the standard input rate.
22K combined input + requested output
Context: 1.1M
Max output: 128K
1M context tokens remain before the published window limit.
Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.
Ranked comparison
Sorted by estimated monthly cost. Select a row to inspect its rates and constraints above. Deprecated and retired models are hidden until you widen the lifecycle filter.
| Model | Per request | Per month | Context fit | Source |
|---|---|---|---|---|
| DeepSeek-V4-Flash | $0.002674 | $58.83 | 2.2% | Official |
| Mistral Small 4 | $0.003525 | $77.55 | 8.6% | Official |
| GPT-5.4-nano | $0.005600 | $123.20 | 5.5% | Official |
| DeepSeek-V4-Pro | $0.008283 | $182.23 | 2.2% | Official |
| Gemini 3.5 Flash-Lite | $0.009650 | $212.30 | 2.1% | Official |
| Mistral Large 3 | $0.0108 | $236.50 | 8.6% | Official |
| Grok Build 0.1 | $0.0200 | $440.00 | 8.6% | Official |
| GPT-5.4-mini | $0.0206 | $453.75 | 5.5% | Official |
| Grok 4.3 | $0.0248 | $544.50 | 2.2% | Official |
| Claude Haiku 4.5 | $0.0255 | $561.00 | 11% | Official |
| GPT-5.6 Luna | $0.0275 | $605.00 | 2.1% | Official |
| Gemini 3.6 Flash | $0.0383 | $841.50 | 2.1% | Official |
| Mistral Medium 3.5 | $0.0383 | $841.50 | 8.6% | Official |
| Gemini 3.5 Flash | $0.0413 | $907.50 | 2.1% | Official |
| Grok 4.5 | $0.0435 | $957.00 | 4.4% | Official |
| Claude Sonnet 5 | $0.0510 | $1,122.00 | 2.2% | Official |
| Gemini 3.1 Pro Preview | $0.0550 | $1,210.00 | 2.1% | Official |
| GPT-5.6 Terra | $0.0688 | $1,512.50 | 2.1% | Official |
| GPT-5.4 | $0.0688 | $1,512.50 | 2.1% | Official |
| Claude Sonnet 4.6 | $0.0765 | $1,683.00 | 2.2% | Official |
| Claude Sonnet 4.5 | $0.0765 | $1,683.00 | 11% | Official |
| Claude Opus 4.8 | $0.1275 | $2,805.00 | 2.2% | Official |
| Claude Opus 4.7 | $0.1275 | $2,805.00 | 2.2% | Official |
| Claude Opus 4.6 | $0.1275 | $2,805.00 | 2.2% | Official |
| Claude Opus 4.5 | $0.1275 | $2,805.00 | 11% | Official |
| GPT-5.6 Sol | $0.1375 | $3,025.00 | 2.1% | Official |
| GPT-5.5 | $0.1375 | $3,025.00 | 2.1% | Official |
| Claude Fable 5 | $0.2550 | $5,610.00 | 2.2% | Official |
| Claude Mythos 5 | $0.2550 | $5,610.00 | 2.2% | Official |
Transparent methodology
Cached tokens are carved out of total input. In Standard mode, only that subset receives a recorded cache rate; Batch uses its own published input rate. Published cache-write fees are disclosed, not silently added.
Input cost and output cost form a per-request estimate. Requests per day multiplied by working days produces the monthly volume, and announced price changes blend into the monthly and yearly projections.
Input plus requested output is compared with the context window, and crossing a documented long-context threshold moves the whole request to the provider's tier rates. Output is also checked independently when the provider publishes a separate maximum.
Common questions
The calculator multiplies the selected model's published per-million-token input and output rates by your tokens per request, then multiplies the request estimate by requests per working day and working days per month. When a provider has announced a dated price change, such as introductory pricing with a published end date, the monthly and yearly projections blend the current rates with the announced successor rates across that horizon; the per-request figure always uses today's rates.
No. Cached input is treated as a subset of the total input. The cached subset receives a published cached-input rate when one is recorded; otherwise it receives the ordinary input rate. Cache reads are not the whole story: several providers also bill to write cache entries, so the calculator shows the model's published cache-write prices beside the savings row without adding them to the estimate. DeepSeek publishes cache-hit and cache-miss input prices rather than a write fee, and the tool labels its savings row accordingly.
No. Batch mode appears only when the model registry contains official Batch input and output rates. The calculator does not assume a percentage discount, combine Batch with a cached-input discount, or infer missing prices. Documented long-context tiers still apply in Batch mode, scaling the published batch rates by the provider's tier multipliers.
No. It covers the recorded text-token prices, documented long-context tiers, and announced successor rates. Taxes, regional pricing, minimum commitments, cache-write and cache-storage fees, searches, grounding, tools, fine-tuning, media, partner platforms, and other feature charges may change the actual bill.