API prices

Every tier the calculator knows about, on one page. Headline rates are what providers publish; the last column is what a request of your shape actually costs on each, which is the only figure that ranks them like for like. Researched to 2026-09-13.

Price them against your own request

Read 98% Written 2% Full rate 0%

98% of input tokens billed as cache hits, 2% as cache writes, 0% at the plain input rate. Measured, not assumed: the token-weighted mean of 64,680 recorded coding-agent requests. Almost every billed input token is a cache read, because a conversation re-reads its whole prefix every turn while writing only that turn’s delta — reads grow with the square of the turn count and writes only linearly. Output is 0.4% of input, so this is a prefill-bound workload, not a decode-bound one.

53 tiers · 15 superseded hidden · 3 too small for this request, sorted last

API price tiers with their per-million rates as actually billed — batch and flex discounts already applied — and the cost of one request at the selected workload shape.
TierScoreInputCached inCache writeOutputContextAt your mix
Xiaomi · MiMo V2.5 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 22 intelligence 22 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.140$0.0028$0$0.2801.05M$0.0039 per M tokens
Meta · Muse Spark 1.2 contributor · Meta trains on your requests trains on requests window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 40 intelligence 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.100$0.0020$0.2001.05M$0.0048 per M tokens 1.2×
Meta · Muse Spark 1.3 contributor · Meta trains on your requests trains on requests window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 48 intelligence 48 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.100$0.0020$0.2001.05M$0.0048 per M tokens 1.2×
Xiaomi · MiMo V2.5 Pro standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 26 intelligence 26 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.435$0.0036$0$0.8701.05M$0.0072 per M tokens 1.8×
DeepSeek · DeepSeek V4.1 Flash peak / off-peak · from 2026-09-10 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 40 intelligence 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-11 (opens in a new tab)$0.300$0.0060$0.300$1.201.05M$0.010 per M tokens 2.6×
OpenAI · GPT-5.6 Luna Flex / Batch · 50% multiplier window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount38 intelligence 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.100$0.010$0.125$0.6001.05M$0.015 per M tokens 3.8×
Google · Gemini 3.1 Flash-Lite Batch / Flex — 50% multiplier window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount$0.125$0.013$0.7501.05M$0.018 per M tokens 4.6×
Alibaba · Qwen3.8-Flash standard · explicit cache window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small $0.150$0.016$0.200$0.4701M$0.022 per M tokens 5.5×
Google · Gemini 3.5 Flash-Lite Batch / Flex window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small $0.150$0.020$1.251.05M$0.028 per M tokens 7.1×
OpenAI · GPT-5.6 Luna standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 38 intelligence 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.200$0.020$0.250$1.201.05M$0.030 per M tokens 7.6×
Z.ai · GLM-5.3-Flash list rate · from 2026-09-10 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 42 intelligence 42 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.150$0.030$0.5001.05M$0.034 per M tokens 8.8×
Google · Gemini 3.1 Flash-Lite standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small $0.250$0.025$1.501.05M$0.036 per M tokens 9.1×
Google · Gemini 3.5 Flash-Lite standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small $0.300$0.030$2.501.05M$0.046 per M tokens 11.7×
DeepSeek · DeepSeek V4 Pro peak / off-peak · from 2026-08-16 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 36 intelligence 36 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.32$0.044$1.32$3.961.05M$0.052 per M tokens 13.3×
Google · Gemini 3.7 Flash Batch / Flex · 50% multiplier · promotional through 2026-12-31 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount39 intelligence 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.375$0.037$1.881.05M$0.052 per M tokens 13.3×
Google · Gemini 3.8 Flash Batch / Flex · 50% multiplier · promotional through 2026-12-31 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount41 intelligence 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.375$0.037$1.881.05M$0.052 per M tokens 13.3×
Google · Gemini 3.7 Flash standard · promotional through 2026-12-31 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 39 intelligence 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.750$0.075$3.751.05M$0.104 per M tokens 26.6×
Google · Gemini 3.7 Flash Batch / Flex · 50% multiplier · list rate from 2027-01-01 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount39 intelligence 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.750$0.075$3.751.05M$0.104 per M tokens 26.6×
Google · Gemini 3.8 Flash standard · promotional through 2026-12-31 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 41 intelligence 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.750$0.075$3.751.05M$0.104 per M tokens 26.6×
Google · Gemini 3.8 Flash Batch / Flex · 50% multiplier · list rate from 2027-01-01 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount41 intelligence 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.750$0.075$3.751.05M$0.104 per M tokens 26.6×
Anthropic · Claude Sonnet 5 Batch API · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount38 intelligence 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.00$0.100$1.25$5.001M$0.144 per M tokens 36.7×
OpenAI · GPT-5.6 Terra Flex / Batch · 50% multiplier window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount42 intelligence 42 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.00$0.100$1.25$6.001.05M$0.148 per M tokens 37.8×
Meta · Muse Spark 1.2 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 40 intelligence 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.25$0.150$4.251.05M$0.189 per M tokens 48.4×
Meta · Muse Spark 1.3 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 48 intelligence 48 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.25$0.150$4.251.05M$0.189 per M tokens 48.4×
Google · Gemini 3.7 Flash list rate · from 2027-01-01 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 39 intelligence 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.50$0.150$7.501.05M$0.208 per M tokens 53.1×
Google · Gemini 3.8 Flash list rate · from 2027-01-01 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 41 intelligence 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.50$0.150$7.501.05M$0.208 per M tokens 53.1×
Anthropic · Claude Sonnet 5 standard · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 38 intelligence 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$2.00$0.200$2.50$10.001M$0.287 per M tokens 73.4×
OpenAI · GPT-5.6 Sol Flex / Batch · 50% multiplier window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount47 intelligence 47 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$2.00$0.200$2.50$10.001.05M$0.287 per M tokens 73.4×
OpenAI · GPT-5.6 Terra standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 42 intelligence 42 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$2.00$0.200$2.50$12.001.05M$0.296 per M tokens 75.5×
Z.ai · GLM-5.3 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 45 intelligence 45 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.40$0.260$4.401.05M$0.300 per M tokens 76.7×
Z.ai · GLM-5.2 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 39 intelligence 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.40$0.260$4.401.05M$0.300 per M tokens 76.7×
Alibaba · Qwen3.8-Max standard · snapshot 0902 · international window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 40 intelligence 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$2.00$0.250$2.50$6.001M$0.319 per M tokens 81.6×
Alibaba · Qwen3.7-Max standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 30 intelligence 30 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$2.50$0.250$3.13$7.501M$0.338 per M tokens 86.3×
Anthropic · Claude Fable 5.1 Batch API · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount53 intelligence 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$5.00$0.125$6.25$25.001M$0.352 per M tokens 89.9×
Anthropic · Claude Opus 5 Batch API · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount51 intelligence 51 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$2.50$0.250$3.13$12.501M$0.359 per M tokens 91.7×
Moonshot AI · Kimi K3 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 44 intelligence 44 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$3.00$0.300$3.00$15.001.05M$0.416 per M tokens 106.3×
Google · Gemini 3.1 Pro Preview · ≤200K preview window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 30 intelligence 30 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$2.00$0.200$12.001.05M$0.546 per M tokens 139.5×
OpenAI · GPT-5.6 Sol standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 47 intelligence 47 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$4.00$0.400$5.00$20.001.05M$0.574 per M tokens 146.8×
Anthropic · Claude Fable 5.1 standard · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 53 intelligence 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$10.00$0.250$12.50$50.001M$0.704 per M tokens 179.8×
Anthropic · Claude Opus 5 standard · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 51 intelligence 51 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$5.00$0.500$6.25$25.001M$0.718 per M tokens 183.5×
Anthropic · Claude Fable 5 Batch API · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount50 intelligence 50 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$5.00$0.500$6.25$25.001M$0.718 per M tokens 183.5×
Anthropic · Claude Mythos 5 Batch API · 5m cache write preview window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount$5.00$0.500$6.25$25.001M$0.718 per M tokens 183.5×
OpenAI · GPT-6 Astra Batch API window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount53 intelligence 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$5.00$0.500$6.25$25.001.05M$0.718 per M tokens 183.5×
xAI · Grok 4.6 standard · auto long-context ≥200K window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 44 intelligence 44 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$2.00$0.500$6.00500K$1.11 per M tokens 282.7×
Alibaba · Qwen3.7-Plus standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 26 intelligence 26 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.400$1.601M$1.22 per M tokens 310.6×
Anthropic · Claude Fable 5 standard · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 50 intelligence 50 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$10.00$1.00$12.50$50.001M$1.44 per M tokens 366.9×
Anthropic · Claude Mythos 5 limited availability · 5m cache write preview window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small $10.00$1.00$12.50$50.001M$1.44 per M tokens 366.9×
Anthropic · Claude Opus 5 fast mode · 5m cache write preview window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small $10.00$1.00$12.50$50.001M$1.44 per M tokens 366.9×
OpenAI · GPT-6 Astra standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 53 intelligence 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$10.00$1.00$12.50$50.001.05M$1.44 per M tokens 366.9×
OpenAI · GPT-5.6 Cyber standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small $12.50$1.25$15.63$75.00400K$1.85 per M tokens 472.1×
Mistral AI · Mistral Medium 3.5 Batch · 50% multiplier window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount15 intelligence 15 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$0.750$0.075$3.75256K$0.104 per M tokens
Anthropic · Claude Haiku 4.5 standard · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 15 intelligence 15 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.00$0.100$1.25$5.00200K$0.144 per M tokens
Mistral AI · Mistral Medium 3.5 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small 15 intelligence 15 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)$1.50$0.150$7.50256K$0.208 per M tokens

How to read this

  • The first six columns are the provider's published rates, per million tokens. A dash means the provider does not publish that line at all — not that it is free. Where no cache-hit rate exists, the calculator charges the ordinary input rate.
  • "At your mix" is the only column that ranks tiers honestly. It prices one request of the shape set above on every tier and converts to dollars per million tokens, so caching, long-context surcharges, batch discounts and time-of-day windows are already inside the number. Change the shape and the order changes with it — which is the point.
  • A capability score beside a price is there to stop a category error. The cheapest row is frequently the cheapest because it answers an easier class of question. Sort by capability per dollar to see the trade rather than one half of it.
  • Tiers with a discount window are priced at the share a continuous service would naturally land there — the share of the week the peak schedule leaves over, not the best case. Days count as well as hours: a peak that runs on weekdays only leaves both weekend days off-peak. A batch workload that can wait for the cheap hours does better; an interactive one does not.
    DeepSeek V4 Flash: peak 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri, so 79% of a round-the-clock workload is costed off-peak.
    DeepSeek V4 Pro: peak 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri, so 79% of a round-the-clock workload is costed off-peak.
    DeepSeek V4.1 Flash: peak 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri, so 79% of a round-the-clock workload is costed off-peak.
  • Retired tiers stay resolvable but are hidden by default, so a shared scenario URL from three months ago still opens. Tick the box to see what a rate replaced.
Inference Economics · static, local-first calculationCatalog 0.8.1 · cutoff 2026-09-13 · app v0.8.1 · build 1c759ca