API prices
Every tier the calculator knows about, on one page. Headline rates are what providers publish; the last column is what a request of your shape actually costs on each, which is the only figure that ranks them like for like. Researched to 2026-09-13.
Price them against your own request
Read 98% Written 2% Full rate 0%
98% of input tokens billed as cache hits, 2% as cache writes, 0% at the plain input rate. Measured, not assumed: the token-weighted mean of 64,680 recorded coding-agent requests. Almost every billed input token is a cache read, because a conversation re-reads its whole prefix every turn while writing only that turn’s delta — reads grow with the square of the turn count and writes only linearly. Output is 0.4% of input, so this is a prefill-bound workload, not a decode-bound one.
| Tier | Score | Input | Cached in | Cache write | Output | Context | At your mix |
|---|---|---|---|---|---|---|---|
| Xiaomi · MiMo V2.5 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 22 intelligence 22 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.140 | $0.0028 | $0 | $0.280 | 1.05M | $0.0039 per M tokens |
| Meta · Muse Spark 1.2 contributor · Meta trains on your requests trains on requests window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 40 intelligence 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.100 | $0.0020 | — | $0.200 | 1.05M | $0.0048 per M tokens 1.2× |
| Meta · Muse Spark 1.3 contributor · Meta trains on your requests trains on requests window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 48 intelligence 48 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.100 | $0.0020 | — | $0.200 | 1.05M | $0.0048 per M tokens 1.2× |
| Xiaomi · MiMo V2.5 Pro standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 26 intelligence 26 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.435 | $0.0036 | $0 | $0.870 | 1.05M | $0.0072 per M tokens 1.8× |
| DeepSeek · DeepSeek V4.1 Flash peak / off-peak · from 2026-09-10 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 40 intelligence 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-11 (opens in a new tab) | $0.300 | $0.0060 | $0.300 | $1.20 | 1.05M | $0.010 per M tokens 2.6× |
| OpenAI · GPT-5.6 Luna Flex / Batch · 50% multiplier window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 38 intelligence 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.100 | $0.010 | $0.125 | $0.600 | 1.05M | $0.015 per M tokens 3.8× |
| Google · Gemini 3.1 Flash-Lite Batch / Flex — 50% multiplier window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | — | $0.125 | $0.013 | — | $0.750 | 1.05M | $0.018 per M tokens 4.6× |
| Alibaba · Qwen3.8-Flash standard · explicit cache window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | — | $0.150 | $0.016 | $0.200 | $0.470 | 1M | $0.022 per M tokens 5.5× |
| Google · Gemini 3.5 Flash-Lite Batch / Flex window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | — | $0.150 | $0.020 | — | $1.25 | 1.05M | $0.028 per M tokens 7.1× |
| OpenAI · GPT-5.6 Luna standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 38 intelligence 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.200 | $0.020 | $0.250 | $1.20 | 1.05M | $0.030 per M tokens 7.6× |
| Z.ai · GLM-5.3-Flash list rate · from 2026-09-10 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 42 intelligence 42 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.150 | $0.030 | — | $0.500 | 1.05M | $0.034 per M tokens 8.8× |
| Google · Gemini 3.1 Flash-Lite standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | — | $0.250 | $0.025 | — | $1.50 | 1.05M | $0.036 per M tokens 9.1× |
| Google · Gemini 3.5 Flash-Lite standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | — | $0.300 | $0.030 | — | $2.50 | 1.05M | $0.046 per M tokens 11.7× |
| DeepSeek · DeepSeek V4 Pro peak / off-peak · from 2026-08-16 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 36 intelligence 36 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.32 | $0.044 | $1.32 | $3.96 | 1.05M | $0.052 per M tokens 13.3× |
| Google · Gemini 3.7 Flash Batch / Flex · 50% multiplier · promotional through 2026-12-31 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 39 intelligence 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.375 | $0.037 | — | $1.88 | 1.05M | $0.052 per M tokens 13.3× |
| Google · Gemini 3.8 Flash Batch / Flex · 50% multiplier · promotional through 2026-12-31 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 41 intelligence 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.375 | $0.037 | — | $1.88 | 1.05M | $0.052 per M tokens 13.3× |
| Google · Gemini 3.7 Flash standard · promotional through 2026-12-31 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 39 intelligence 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.750 | $0.075 | — | $3.75 | 1.05M | $0.104 per M tokens 26.6× |
| Google · Gemini 3.7 Flash Batch / Flex · 50% multiplier · list rate from 2027-01-01 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 39 intelligence 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.750 | $0.075 | — | $3.75 | 1.05M | $0.104 per M tokens 26.6× |
| Google · Gemini 3.8 Flash standard · promotional through 2026-12-31 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 41 intelligence 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.750 | $0.075 | — | $3.75 | 1.05M | $0.104 per M tokens 26.6× |
| Google · Gemini 3.8 Flash Batch / Flex · 50% multiplier · list rate from 2027-01-01 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 41 intelligence 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.750 | $0.075 | — | $3.75 | 1.05M | $0.104 per M tokens 26.6× |
| Anthropic · Claude Sonnet 5 Batch API · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 38 intelligence 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.00 | $0.100 | $1.25 | $5.00 | 1M | $0.144 per M tokens 36.7× |
| OpenAI · GPT-5.6 Terra Flex / Batch · 50% multiplier window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 42 intelligence 42 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.00 | $0.100 | $1.25 | $6.00 | 1.05M | $0.148 per M tokens 37.8× |
| Meta · Muse Spark 1.2 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 40 intelligence 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.25 | $0.150 | — | $4.25 | 1.05M | $0.189 per M tokens 48.4× |
| Meta · Muse Spark 1.3 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 48 intelligence 48 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.25 | $0.150 | — | $4.25 | 1.05M | $0.189 per M tokens 48.4× |
| Google · Gemini 3.7 Flash list rate · from 2027-01-01 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 39 intelligence 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.50 | $0.150 | — | $7.50 | 1.05M | $0.208 per M tokens 53.1× |
| Google · Gemini 3.8 Flash list rate · from 2027-01-01 window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 41 intelligence 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.50 | $0.150 | — | $7.50 | 1.05M | $0.208 per M tokens 53.1× |
| Anthropic · Claude Sonnet 5 standard · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 38 intelligence 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $2.00 | $0.200 | $2.50 | $10.00 | 1M | $0.287 per M tokens 73.4× |
| OpenAI · GPT-5.6 Sol Flex / Batch · 50% multiplier window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 47 intelligence 47 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $2.00 | $0.200 | $2.50 | $10.00 | 1.05M | $0.287 per M tokens 73.4× |
| OpenAI · GPT-5.6 Terra standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 42 intelligence 42 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $2.00 | $0.200 | $2.50 | $12.00 | 1.05M | $0.296 per M tokens 75.5× |
| Z.ai · GLM-5.3 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 45 intelligence 45 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.40 | $0.260 | — | $4.40 | 1.05M | $0.300 per M tokens 76.7× |
| Z.ai · GLM-5.2 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 39 intelligence 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.40 | $0.260 | — | $4.40 | 1.05M | $0.300 per M tokens 76.7× |
| Alibaba · Qwen3.8-Max standard · snapshot 0902 · international window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 40 intelligence 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $2.00 | $0.250 | $2.50 | $6.00 | 1M | $0.319 per M tokens 81.6× |
| Alibaba · Qwen3.7-Max standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 30 intelligence 30 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $2.50 | $0.250 | $3.13 | $7.50 | 1M | $0.338 per M tokens 86.3× |
| Anthropic · Claude Fable 5.1 Batch API · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 53 intelligence 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $5.00 | $0.125 | $6.25 | $25.00 | 1M | $0.352 per M tokens 89.9× |
| Anthropic · Claude Opus 5 Batch API · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 51 intelligence 51 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $2.50 | $0.250 | $3.13 | $12.50 | 1M | $0.359 per M tokens 91.7× |
| Moonshot AI · Kimi K3 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 44 intelligence 44 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $3.00 | $0.300 | $3.00 | $15.00 | 1.05M | $0.416 per M tokens 106.3× |
| Google · Gemini 3.1 Pro Preview · ≤200K preview window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 30 intelligence 30 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $2.00 | $0.200 | — | $12.00 | 1.05M | $0.546 per M tokens 139.5× |
| OpenAI · GPT-5.6 Sol standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 47 intelligence 47 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $4.00 | $0.400 | $5.00 | $20.00 | 1.05M | $0.574 per M tokens 146.8× |
| Anthropic · Claude Fable 5.1 standard · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 53 intelligence 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $10.00 | $0.250 | $12.50 | $50.00 | 1M | $0.704 per M tokens 179.8× |
| Anthropic · Claude Opus 5 standard · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 51 intelligence 51 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $5.00 | $0.500 | $6.25 | $25.00 | 1M | $0.718 per M tokens 183.5× |
| Anthropic · Claude Fable 5 Batch API · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 50 intelligence 50 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $5.00 | $0.500 | $6.25 | $25.00 | 1M | $0.718 per M tokens 183.5× |
| Anthropic · Claude Mythos 5 Batch API · 5m cache write preview window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | — | $5.00 | $0.500 | $6.25 | $25.00 | 1M | $0.718 per M tokens 183.5× |
| OpenAI · GPT-6 Astra Batch API window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 53 intelligence 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $5.00 | $0.500 | $6.25 | $25.00 | 1.05M | $0.718 per M tokens 183.5× |
| xAI · Grok 4.6 standard · auto long-context ≥200K window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 44 intelligence 44 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $2.00 | $0.500 | — | $6.00 | 500K | $1.11 per M tokens 282.7× |
| Alibaba · Qwen3.7-Plus standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 26 intelligence 26 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.400 | — | — | $1.60 | 1M | $1.22 per M tokens 310.6× |
| Anthropic · Claude Fable 5 standard · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 50 intelligence 50 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $10.00 | $1.00 | $12.50 | $50.00 | 1M | $1.44 per M tokens 366.9× |
| Anthropic · Claude Mythos 5 limited availability · 5m cache write preview window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | — | $10.00 | $1.00 | $12.50 | $50.00 | 1M | $1.44 per M tokens 366.9× |
| Anthropic · Claude Opus 5 fast mode · 5m cache write preview window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | — | $10.00 | $1.00 | $12.50 | $50.00 | 1M | $1.44 per M tokens 366.9× |
| OpenAI · GPT-6 Astra standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 53 intelligence 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $10.00 | $1.00 | $12.50 | $50.00 | 1.05M | $1.44 per M tokens 366.9× |
| OpenAI · GPT-5.6 Cyber standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | — | $12.50 | $1.25 | $15.63 | $75.00 | 400K | $1.85 per M tokens 472.1× |
| Mistral AI · Mistral Medium 3.5 Batch · 50% multiplier window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small rates shown after the 50% batch discount | 15 intelligence 15 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $0.750 | $0.075 | — | $3.75 | 256K | $0.104 per M tokens |
| Anthropic · Claude Haiku 4.5 standard · 5m cache write window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 15 intelligence 15 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.00 | $0.100 | $1.25 | $5.00 | 200K | $0.144 per M tokens |
| Mistral AI · Mistral Medium 3.5 standard window too small for this requestinput limit too small for this requestoutput cap too small for this requestoutput cap may be too small | 15 intelligence 15 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab) | $1.50 | $0.150 | — | $7.50 | 256K | $0.208 per M tokens |
How to read this
- The first six columns are the provider's published rates, per million tokens. A dash means the provider does not publish that line at all — not that it is free. Where no cache-hit rate exists, the calculator charges the ordinary input rate.
- "At your mix" is the only column that ranks tiers honestly. It prices one request of the shape set above on every tier and converts to dollars per million tokens, so caching, long-context surcharges, batch discounts and time-of-day windows are already inside the number. Change the shape and the order changes with it — which is the point.
- A capability score beside a price is there to stop a category error. The cheapest row is frequently the cheapest because it answers an easier class of question. Sort by capability per dollar to see the trade rather than one half of it.
- Tiers with a discount window are priced at the share a continuous service would naturally
land there — the share of the week the peak schedule leaves over, not the best case. Days count as well as
hours: a peak that runs on weekdays only leaves both weekend days off-peak. A batch workload that
can wait for the cheap hours does better; an interactive one does not.
DeepSeek V4 Flash: peak 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri, so 79% of a round-the-clock workload is costed off-peak.
DeepSeek V4 Pro: peak 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri, so 79% of a round-the-clock workload is costed off-peak.
DeepSeek V4.1 Flash: peak 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri, so 79% of a round-the-clock workload is costed off-peak. - Retired tiers stay resolvable but are hidden by default, so a shared scenario URL from three months ago still opens. Tick the box to see what a rate replaced.