Source-backed AI Wiki tool

AI API cost and context planner

Turn a token workload into per-request, monthly, and annual estimates. Compare official prices and see which model context windows can actually hold the request.

No API keyInputs stay in your browserFacts reviewed 2026-07-23
This is a planning estimate, not a quote or invoice. Review the selected model's effective dates and caveats, then confirm the current provider price before committing a budget.

Private, browser-only estimate

No API key is needed. Your workload inputs stay on this device and are not saved or sent to a provider. The address bar carries your settings so the estimate can be shared.

Workload

Tokens and request volume

Cached input is a subset of input, never an additional token count.

Presets (illustrative defaults, not provider guidance)

Pricing mode

A published cached-input rate is applied only to the cached subset. Otherwise those tokens use the regular input rate.

Selected estimate

GPT-5.6 Terra

gpt-5.6-terra

Per request

$0.0688

22K total tokens

Per month

$1,512.50

22K requests

Per year

$18,150.00

12 equal months

Cost composition

Input
$0.0388
Output
$0.0300
Cache read savings
-$0.0113

Cache writes bill at $3.125 per 1M and are not modeled here. OpenAI bills prompt-cache writes for this model at 1.25x the standard input rate.

Published rates

Input / 1M
$2.50
Output / 1M
$15.00
Cached input / 1M
$0.2500

Context and output fit

22K combined input + requested output

2.1% used

Context: 1.1M

Max output: 128K

1M context tokens remain before the published window limit.

ActiveRates effective Jul 23, 2026

Standard short-context processing. Requests above 272K input tokens are billed at 2x input and 1.5x output for the full request.

Ranked comparison

Compare sourced API prices

Sorted by estimated monthly cost. Select a row to inspect its rates and constraints above. Deprecated and retired models are hidden until you widen the lifecycle filter.

API model costs for the entered workload, ranked by monthly estimate
ModelPer requestPer monthContext fitSource
DeepSeek-V4-Flash$0.002674$58.832.2%Official
Mistral Small 4$0.003525$77.558.6%Official
GPT-5.4-nano$0.005600$123.205.5%Official
DeepSeek-V4-Pro$0.008283$182.232.2%Official
Gemini 3.5 Flash-Lite$0.009650$212.302.1%Official
Mistral Large 3$0.0108$236.508.6%Official
Grok Build 0.1$0.0200$440.008.6%Official
GPT-5.4-mini$0.0206$453.755.5%Official
Grok 4.3$0.0248$544.502.2%Official
Claude Haiku 4.5$0.0255$561.0011%Official
GPT-5.6 Luna$0.0275$605.002.1%Official
Gemini 3.6 Flash$0.0383$841.502.1%Official
Mistral Medium 3.5$0.0383$841.508.6%Official
Gemini 3.5 Flash$0.0413$907.502.1%Official
Grok 4.5$0.0435$957.004.4%Official
Claude Sonnet 5$0.0510$1,122.002.2%Official
Gemini 3.1 Pro Preview$0.0550$1,210.002.1%Official
GPT-5.6 Terra$0.0688$1,512.502.1%Official
GPT-5.4$0.0688$1,512.502.1%Official
Claude Sonnet 4.6$0.0765$1,683.002.2%Official
Claude Sonnet 4.5$0.0765$1,683.0011%Official
Claude Opus 4.8$0.1275$2,805.002.2%Official
Claude Opus 4.7$0.1275$2,805.002.2%Official
Claude Opus 4.6$0.1275$2,805.002.2%Official
Claude Opus 4.5$0.1275$2,805.0011%Official
GPT-5.6 Sol$0.1375$3,025.002.1%Official
GPT-5.5$0.1375$3,025.002.1%Official
Claude Fable 5$0.2550$5,610.002.2%Official
Claude Mythos 5$0.2550$5,610.002.2%Official
Model facts reviewed Jul 23, 2026. When input crosses a documented long-context threshold, the whole request is billed at the tier rates, in Batch mode too (batch rates scale by the provider's tier multipliers). Announced successor rates blend into the monthly and yearly projections. Estimates exclude taxes, minimum commitments, cache-write and cache-storage fees, web search, tools, grounding, fine-tuning, and provider-specific regional or partner-platform charges.

Transparent methodology

What the estimate includes

01

Split input correctly

Cached tokens are carved out of total input. In Standard mode, only that subset receives a recorded cache rate; Batch uses its own published input rate. Published cache-write fees are disclosed, not silently added.

02

Scale request volume

Input cost and output cost form a per-request estimate. Requests per day multiplied by working days produces the monthly volume, and announced price changes blend into the monthly and yearly projections.

03

Check the constraints

Input plus requested output is compared with the context window, and crossing a documented long-context threshold moves the whole request to the provider's tier rates. Output is also checked independently when the provider publishes a separate maximum.

Common questions

AI API pricing FAQ

How is monthly AI API cost calculated?

The calculator multiplies the selected model's published per-million-token input and output rates by your tokens per request, then multiplies the request estimate by requests per working day and working days per month. When a provider has announced a dated price change, such as introductory pricing with a published end date, the monthly and yearly projections blend the current rates with the announced successor rates across that horizon; the per-request figure always uses today's rates.

Are cached input tokens added to regular input tokens?

No. Cached input is treated as a subset of the total input. The cached subset receives a published cached-input rate when one is recorded; otherwise it receives the ordinary input rate. Cache reads are not the whole story: several providers also bill to write cache entries, so the calculator shows the model's published cache-write prices beside the savings row without adding them to the estimate. DeepSeek publishes cache-hit and cache-miss input prices rather than a write fee, and the tool labels its savings row accordingly.

Does Batch always cost half as much?

No. Batch mode appears only when the model registry contains official Batch input and output rates. The calculator does not assume a percentage discount, combine Batch with a cached-input discount, or infer missing prices. Documented long-context tiers still apply in Batch mode, scaling the published batch rates by the provider's tier multipliers.

Does the estimate include every provider fee?

No. It covers the recorded text-token prices, documented long-context tiers, and announced successor rates. Taxes, regional pricing, minimum commitments, cache-write and cache-storage fees, searches, grounding, tools, fine-tuning, media, partner platforms, and other feature charges may change the actual bill.