Open reference · v0.8.1

Price the decision, expose the uncertainty.

A local inference system only wins under a workload, a utilization rate, and a performance profile. Explore the crossover point against hosted API spend — with every assumption in view.

01 · model a scenario

Calculator

Start from a curated scenario or build your own. Results recompute locally, are shareable as a URL, and can be saved in this browser.

A team runs long-context coding agents all day. Does a single dense node beat frontier API spend?

Break-even 7moAt 3y $1.7M Live

Assumptions

Selection
8 × B300 · 2,304 GB memory
moe · 2,800B params · 1,048,576 context
mixed precision
$5.00 in · $25.00 out per M tokens
The ordinary synchronous rate.
Every published rate in this catalog is the provider’s default, unpinned routing.

What these lists include, and at which price

Throughput evidence
SGLang v0.5.18 @ 71de97b2, TP8 + DCP8, native MXFP4 checkpoint
estimated medium confidence
Decode 1,240 tok/s
Prefill 9,920 tok/s
Split at 8,192 input / 1,024 output tokens per request. On ratio alone, request throughput is understated here; see the caveats.
The share of the rate above that your own deployment reaches.

The profile supplies both rates; achievable throughput scales them. Nothing here was measured on your hardware.

Workload
Prompt length for one typical request, including anything served from cache. Drives prefill work locally and the input rate on the API side.
Tokens generated per request. These are the expensive ones: decode is one token at a time, and every API tier prices output above input.
Share of every input token you send that the provider already holds.
Shares one 100% budget with cache reads; 30% is left to divide.
Read 70% Written 5% Full rate 25%
Cache scales with this and with context, so it usually sets the hardware — not the weights.
Tied to input plus output (25,000). This model's window holds 1,048,576.
What the reference deployment stores. An fp8 path exists and would halve the part of the cache it reaches: a recurrent layer stores its state at fp32 and its convolution window at bf16 whatever this says.
Here, tensor parallelism holds 5.2 times as much cache as context parallelism.
Operations
The share of the day the hardware is serving requests. Not how fast it serves them — that is Achievable throughput above, and setting both to the same figure derates the case twice.
What you pay per kilowatt-hour at the meter. Multiplied by PUE below, so this is the IT draw rather than the facility's.
Power usage effectiveness: total facility draw per watt drawn by the equipment. 1.3 means cooling and distribution add 30%.
Yearly support and maintenance, as a share of the all-in acquisition cost.
Commercial & quality
Published scores put the local model at 0.86× the API one. Yours to agree with.

Live result

Recomputed locally · no data leaves this page
Break-even duration
7mo
from acquisition
Net position at horizon
$1,667,534
after $420,000 all-in acquisition
Effective cost / M tokens
$1.79
API equivalent $6.53
Request throughput
535.7 / hr
decode-bound · 10,713,600 input tokens / hr
Break-even utilization
16%
you assumed 60%
Throughput evidence
medium
taken at face value
estimated

At 60% utilization DGX B300 costs less than Anthropic Claude Opus 5 after 7mo, and is $1,667,534 ahead by the end of the 3-year horizon — if your deployment reaches the modeled throughput, which is estimated evidence at medium confidence.

Read this result with 3 caveats

  • This request's 20,000 input and 5,000 output tokens are not in the ratio the "base" profile's rates were split at, 8,192 to 1,024.

    Request rate is the lower of decode ÷ output tokens and prefill ÷ input tokens, and the two limits meet only at the ratio the rates were split at. Here decode runs out first: 1,240 tokens/s ÷ 5,000 output tokens admits 893 requests an hour at full utilization, where prefill's 9,920 ÷ 20,000 admits 1,786, so 50% of the prefill capacity goes unused. A deployment built for this request would move GPU time from prefill to decode, and at the runs' cost per token it would serve between those two figures; the result is costed at the lower one, so on ratio alone it understates. At 25,000 tokens the request is also longer than the runs' 9,216, and every token of a longer request costs more in both phases, which the calculator does not model and which overstates.

    Rates measured at this request's shape can replace the profile's under Throughput evidence.

  • Aggregate memory fits, but the topology is below official guidance.

    Moonshot AI recommends at least 64 accelerators for Kimi K3; this scenario uses 8 × B300. A capacity check is not a statement that the deployment works or performs.

  • This scenario credits Kimi K3 with more value than Artificial Analysis Intelligence Index v4.3 supports.

    Artificial Analysis Intelligence Index v4.3 puts Kimi K3 at 44 and Claude Opus 5 at 51 on the same scale, a ratio of 0.86. Quality parity is set to 1.00×, so the comparison treats one unit of Kimi K3 output as worth more than that scale suggests. Cost per token is only meaningful between models that can do the same job: a composite index is a blunt instrument, but a gap is worth pricing deliberately rather than by default.

    Set quality parity to what the two models are worth for your workload — the index is a starting point, not an answer.

Cumulative cost curve

  • Self-hosted
  • API
Cumulative self-hosted versus API costA line chart comparing cumulative modeled cost over the selected horizon. A dashed marker shows the modeled break-even month.Total spend to dateMonths from purchase$0$574.5k$1.1M$1.7M$2.3Mbreak-even06 mo1 yr18 mo2 yr30 mo3 yr
showing the whole horizon · drag across the chart to zoom showing month 0 to 36 of 36

At month 36, self-hosted cost is $630,645 and API equivalent is $2,298,180. Break-even lands after 7mo, 20% of the way into the modeled horizon.

What the same workload costs on every current tier

Cost per request at 20,000 input and 5,000 output tokens, with the current cache ratios and routing. Token counts are not comparable across tokenizers, so treat this as a spread, not a ranking of value.
TierPer requestPer 1k requestsvs cheapestScore
Meta · Muse Spark 1.2 contributor · Meta trains on your requests $0.0016$1.6340 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Meta · Muse Spark 1.3 contributor · Meta trains on your requests $0.0016$1.6348 48 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Xiaomi · MiMo V2.5 standard $0.0021$2.141.3×22 22 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Alibaba · Qwen3.8-Flash standard · explicit cache $0.0035$3.522.2×
OpenAI · GPT-5.6 Luna Flex / Batch · 50% multiplier $0.0038$3.772.3×38 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Z.ai · GLM-5.3-Flash list rate · from 2026-09-10 $0.0038$3.822.3×42 42 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.1 Flash-Lite Batch / Flex — 50% multiplier $0.0047$4.682.9×
DeepSeek · DeepSeek V4.1 Flash peak / off-peak · from 2026-09-10 $0.0048$4.762.9×40 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-11 (opens in a new tab)
Xiaomi · MiMo V2.5 Pro standard $0.0066$6.5826 26 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.5 Flash-Lite Batch / Flex $0.0074$7.434.6×
OpenAI · GPT-5.6 Luna standard $0.0075$7.534.6×38 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.1 Flash-Lite standard $0.0094$9.355.7×
Google · Gemini 3.7 Flash Batch / Flex · 50% multiplier · promotional through 2026-12-31 $0.0122$12.157.5×39 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.8 Flash Batch / Flex · 50% multiplier · promotional through 2026-12-31 $0.0122$12.157.5×41 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.5 Flash-Lite standard $0.0147$14.72
Alibaba · Qwen3.7-Plus standard $0.0160$16.009.8×26 26 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
DeepSeek · DeepSeek V4 Pro peak / off-peak · from 2026-08-16 $0.0171$17.1210.5×36 36 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.7 Flash standard · promotional through 2026-12-31 $0.0243$24.3014.9×39 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.7 Flash Batch / Flex · 50% multiplier · list rate from 2027-01-01 $0.0243$24.3014.9×39 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Mistral AI · Mistral Medium 3.5 Batch · 50% multiplier $0.0243$24.3014.9×15 15 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.8 Flash standard · promotional through 2026-12-31 $0.0243$24.3014.9×41 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.8 Flash Batch / Flex · 50% multiplier · list rate from 2027-01-01 $0.0243$24.3014.9×41 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Meta · Muse Spark 1.2 standard $0.0309$30.8518.9×40 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Meta · Muse Spark 1.3 standard $0.0309$30.8518.9×48 48 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Sonnet 5 Batch API · 5m cache write $0.0327$32.6520.1×38 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Haiku 4.5 standard · 5m cache write $0.0327$32.6520.1×15 15 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Z.ai · GLM-5.3 standard $0.0340$34.0420.9×45 45 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Z.ai · GLM-5.2 standard $0.0340$34.0420.9×39 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
OpenAI · GPT-5.6 Terra Flex / Batch · 50% multiplier $0.0377$37.6523.1×42 42 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Alibaba · Qwen3.8-Max standard · snapshot 0902 · international $0.0460$46.0028.3×40 40 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.7 Flash list rate · from 2027-01-01 $0.0486$48.6029.9×39 39 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Mistral AI · Mistral Medium 3.5 standard $0.0486$48.6029.9×15 15 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.8 Flash list rate · from 2027-01-01 $0.0486$48.6029.9×41 41 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
xAI · Grok 4.6 standard · auto long-context ≥200K $0.0490$49.0030.1×44 44 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Alibaba · Qwen3.7-Max standard $0.0566$56.6334.8×30 30 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Sonnet 5 standard · 5m cache write $0.0653$65.3040.1×38 38 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
OpenAI · GPT-5.6 Sol Flex / Batch · 50% multiplier $0.0653$65.3040.1×47 47 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Google · Gemini 3.1 Pro Preview · ≤200K preview $0.0748$74.8045.9×30 30 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
OpenAI · GPT-5.6 Terra standard $0.0753$75.3046.3×42 42 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Opus 5 Batch API · 5m cache write $0.0816$81.6350.1×51 51 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Moonshot AI · Kimi K3 standard $0.0972$97.2059.7×44 44 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
OpenAI · GPT-5.6 Sol standard $0.1306$130.6080.2×47 47 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Fable 5.1 Batch API · 5m cache write $0.1580$158.0097.1×53 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Opus 5 standard · 5m cache write $0.1633$163.25100.3×51 51 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Fable 5 Batch API · 5m cache write $0.1633$163.25100.3×50 50 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Mythos 5 Batch API · 5m cache write preview $0.1633$163.25100.3×
OpenAI · GPT-6 Astra Batch API $0.1633$163.25100.3×53 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Fable 5.1 standard · 5m cache write $0.3160$316.00194.1×53 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Fable 5 standard · 5m cache write $0.3265$326.50200.6×50 50 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
Anthropic · Claude Mythos 5 limited availability · 5m cache write preview $0.3265$326.50200.6×
Anthropic · Claude Opus 5 fast mode · 5m cache write preview $0.3265$326.50200.6×
OpenAI · GPT-6 Astra standard $0.3265$326.50200.6×53 53 of 100 on Artificial Analysis Intelligence Index v4.3, observed 2026-09-08 (opens in a new tab)
OpenAI · GPT-5.6 Cyber standard $0.4706$470.63289.1×

What would have to be true

Solved thresholds, not sampled points: the value of each assumption at which break-even stops landing inside the 3-year horizon.
Utilization must stay above16%assumed 60%
Achievable throughput must stay above27%assumed 100%
Quality parity must stay above0.27×assumed 1.00×

What moves the answer

Each driver varied on its own, with every other assumption held constant. Break-even and API value count the API side at quality parity, the scenario's or the row's own. In cash counts the API's own bill, $6.53 per million tokens in every row, and is left blank in rows at 1.00×, where it would repeat Break-even.
DriverBreak-evenIn cashSelf-host $/MAPI value $/M
Utilization assumed 60%
25%1y 8mo$4.30$6.53
50%9mo$2.15$6.53
75%6mo$1.43$6.53
100%4mo$1.08$6.53
Achievable throughput assumed 100%
50%1y 4mo$3.58$6.53
70%11mo$2.56$6.53
85%9mo$2.11$6.53
100%7mo$1.79$6.53
Quality parity assumed 1.00×
0.75×10mo7mo$1.79$4.90
1.00×7mo$1.79$6.53
1.25×6mo7mo$1.79$8.16

Where the memory goes

Aggregate across 8 accelerators, holding 48 sequences of 25,000 tokens each.
Weights1,561 GB
Runtime overhead312 GB
KV cache55 GB
Free376 GB
Sequences that fit377
Context that fits 308,489

Sequences and context that fit are counted with one recurrent state per sequence, where vLLM keeps two or three per request by default and SGLang admits one request per five state slots, so on either default they are the most that could fit.

Is the cited rate reachable here

A check on the cited rate, not a forecast for your deployment: 1,240 tokens/s against 64 TB/s of published bandwidth and a 1,561 GB checkpoint.
Requests in flight the rate needs 2
The fewest it could need 2
Requests this system can hold377
Cache reads alone 1.4 TB/s
Memory bandwidth available64 TB/s

A decode step reads the weights its requests need, then some of every in-flight request's cache. The weight read is shared by every request in the step, so a high token rate comes from spreading it across many requests at once, or from emitting several tokens a step. The first two rows are that arithmetic under two sets of assumptions about how much a step reads; neither describes a particular serving stack.

This system holds the requests the estimate needs, so the rate and this request size are consistent. Requests this system can hold are counted with one recurrent state each, where vLLM and SGLang both keep more by default, so the system may hold fewer. See how this is derived.

Hourly economics

At the selected workload and utilization.
API-equivalent value / hour $87.45
Raw API spend / hour$87.45
Electricity / hour$2.26
Support / hour$5.75
Net savings / hour$79.43
02 · inspect the inputs

Catalog

Every value is versioned, dated, and tagged with the kind of evidence behind it.

21hardware systems
34open models
68API price tiers
37benchmark observations
194sources

Browse the full catalog with specifications, prices, evidence badges and source links — researched to 2026-09-13.

03 · read the frame

Methodology

Reproducible formulas, explicit exclusions, and a clear quality caveat.

What is calculated

Per-request input, cache-hit, cache-write, and output tokens are priced separately. Throughput is constrained by both decode and prefill; facility electricity applies IT power × PUE. Support is an annual ratio of all-in acquisition cost. Break-even exists only when the hourly API-equivalent value, the API bill for the same tokens times quality parity, exceeds self-hosted variable cost.

Memory, including KV cache

Weights fitting is necessary, not sufficient. The memory check covers weights plus a runtime overhead ratio and the KV cache, which grows with context length and with how many requests are in flight — so a model can pass the capacity check and still be unable to hold more than a couple of concurrent sequences. KV cost is summed over layer groups rather than assumed uniform, because current flagships mix attention types: a sliding-window layer's cache stops growing at the window and a linear-attention layer's state never grows at all. One consequence is worth stating plainly — a compressed latent cache is a single key/value head, so ordinary tensor parallelism gives every accelerator a full copy of it. The weights divide across the system and the cache does not. Models with no published cache profile say so rather than guessing.

Evidence, ranked

Values carry the kind of evidence behind them, strongest first. Measured means a run happened and its result was recorded: MLPerf's closed division, a third-party harness, or logged traffic. It does not mean independent — that is a property of the source, and the Evidence page names the publisher beside every claim and grades each source's independence in the registry above them. One of the sources graded this way is this repository owner's own recorded agent traffic. Vendor-reported is a real measurement the vendor published for its own product — usually on a stack it tuned, so it is an upper bound rather than a promise. Official claim is a published specification or list price: a design fact, not a run of a workload. Estimated is derived here from other published values, and assumption is a planning number this repository chose. Where better public evidence exists, it replaces a vendor's own figure.

Achievable throughput

Every cost per token here scales inversely with the throughput you actually reach, and published throughput figures are tuned figures. The common worry is that open-source serving stacks lag a vendor's own; the best same-hardware evidence does not support it — in MLPerf Inference v6.0 a tuned vLLM submission led its 8× B200 category outright, and retuning that same engine moved its own result by about a fifth. What the data does show is a gap between a tuned deployment and an untuned one. The achievable-throughput control models that gap explicitly, and is kept separate from utilization, which is only how busy the hardware is kept.

What the capability score is, and is not

Scores shown beside models are the Artificial Analysis Intelligence Index, cited with a link to the model's own page there. They are that organisation's measurement, not this repository's, and they carry real limits. A composite collapses nine dissimilar evaluations — agentic tool use, terminal coding, PhD-level science questions, long-context reasoning, hallucination rate — into one number, so two models can tie on the headline and differ sharply at the thing you actually need. The index is also a moving target: it has been revised roughly fourteen times since early 2024, and a revision can change which benchmarks count, how they are weighted, or even which model grades the open-ended answers, so figures from different versions are not strictly comparable. And a score belongs to one specific configuration — the same Claude Sonnet 4.6 checkpoint scores 37 or 48 depending only on whether adaptive reasoning is on — while a self-hosted deployment can differ in quantization, reasoning effort and context length from whatever was benchmarked.

Colour bands compare each score to the best in this catalog rather than to the index's nominal ceiling of 100, because today's frontier sits in the low sixties. The score is here for one reason: every headline on this page is a cost ratio, and a cost ratio between two models of very different capability is a category error wearing the clothes of a finding.

Quality-parity caveat

The quality-parity multiplier scales the value of successful local output against the API comparator. It does not claim the models are equivalent, and it never changes physical throughput.

Exclusions and semantics

Financing, taxes, staffing, downtime, resale, opportunity cost, and production support beyond the configured ratio are excluded. Tokenizers also differ per provider and even per model generation — Anthropic documents that its newest generation produces roughly 30% more tokens for the same text than Sonnet 4.6 and earlier — so equal token counts across API tiers do not describe equal text volumes; affected tiers carry a note. Superseded catalog entries stay available so older shared scenarios keep resolving, but are hidden until you ask for them. The research cutoff is the catalog's last-updated boundary; changing the URL scenario changes assumptions, not the evidence. The repository's docs/architecture.md and docs/PROJECT.md record the full scope.

Independent work this relies on

Two outside efforts do work this repository depends on and has not reproduced. Artificial Analysis produces every capability score shown beside a model or a tier; each one links to the page it was read from, and what a composite index can and cannot tell you is set out above.

Sebastian Raschka publishes per-release architecture notes and an LLM architecture gallery that read new releases down to their attention geometry. Both uses of it here are worth naming. The first is corroboration: his read of GLM-5.3-Flash found the same 34 linear-attention layers against 11 sparse-attention layers this catalog derived from the release's own configuration file, from the same release and with no contact between the two.

The second is disagreement, which was more useful. His published cache-per-token figures match this catalog exactly on Muse Glimmer 30B at 52 KiB and Qwen3.6 27B at 64 KiB, and differ on Gemma 4 31B: 840 KiB against 880 KiB here. The entire 40 KiB gap is the ten global layers, which his figure sizes at half what this catalog does. Two different readings of Gemma 4's configuration land on that half — using the model-wide head dimension where the global layers declare their own larger one, or treating its attention_k_eq_v flag as storing keys and values in one tensor — and this catalog carried the second until 2026-08-21. Which of the two is behind the published figure cannot be told from the figure, so the entry says so rather than picking one. Both traps are now written out in the entry itself.

For the explanations behind what this page only models, his KV cache walkthrough and visual guide to attention variants cover the mechanisms the memory page here turns into numbers. These links open in a new tab.

Inference Economics · static, local-first calculationCatalog 0.8.1 · cutoff 2026-09-13 · app v0.8.1 · build 1c759ca