Evidence

Measured observations stay separate from modeled profiles. Follow the source registry for context.

Benchmark observations

8×H200 · Llama 2 70B

27,417.3 tok/s, counting generated tokens · MLPerf Inference v5.0 · server

measured high confidence

Server result; benchmark harness and latency constraints apply.

8×MI325X · Llama 2 70B

30,724.5 tok/s, counting generated tokens · MLPerf Inference v5.0 · server

measured high confidence

Server result; benchmark harness and latency constraints apply.

DGX B200 · Llama 2 70B

92,338.7 tok/s, counting generated tokens · MLPerf Inference v5.0 preview · server

measured medium confidence

Preview result; benchmark harness and latency constraints apply.

8×H200 · Mixtral 8x7B

57,177 tok/s, counting generated tokens · MLPerf Inference v4.1 · server

measured high confidence
8×H100 · Mixtral 8x7B

50,796 tok/s, counting generated tokens · MLPerf Inference v4.1 · server

measured high confidence
MI300X server · Llama 2 70B

22,021 tok/s, counting generated tokens · MLPerf Inference v4.1 (ROCm) · server

measured high confidence

MLPerf Inference v4.1 closed-division Server result, quoted on AMD's page: AMD's own 8×MI300X submission with EPYC Turin logs 22,020.92 completed tokens per second, against NVIDIA's DGX H100 at 21,605.79. CPU, runtime and latency caveats apply.

H100 server · Llama 2 70B

21,605 tok/s, counting generated tokens · MLPerf Inference v4.1 (TensorRT) · server

measured high confidence

NVIDIA's own MLPerf Inference v4.1 closed-division Server result for DGX H100, 21,605.79 completed tokens per second in its submission log, which AMD's page quotes as the comparison for its MI300X row. CORRECTED 2026-09-13: this note called the figure AMD's report, made without independent audit; it is NVIDIA's own submission, which AMD only quotes.

MI300X server · Llama 2 70B

24,110 tok/s, counting generated tokens · MLPerf Inference v4.1 (ROCm) · offline

measured high confidence

AMD's own MLPerf Inference v4.1 closed-division Offline result for 8×MI300X with EPYC Turin, 24,109.8 tokens per second in its submission log, against NVIDIA's DGX H100 at 24,524.9. CORRECTED 2026-09-13: this row was rocm-mi300x-llama31-405b, labelled Llama 3.1 405B and linked to that model, because AMD's page labels the table row that way. Round 4.1 had no Llama 3.1 405B benchmark, and both figures are the Llama 2 70B Offline results in the two submissions' own logs. On the throughput page the row sat beside an audited 405B result some fifteen times lower per accelerator.

H100 server · Llama 2 70B

24,525 tok/s, counting generated tokens · MLPerf Inference v4.1 (TensorRT) · offline

measured high confidence

NVIDIA's own MLPerf Inference v4.1 closed-division Offline result for DGX H100, 24,524.9 tokens per second in its submission log, quoted on AMD's page as the comparison. CORRECTED 2026-09-13: this row was rocm-h100-llama31-405b, labelled Llama 3.1 405B and linked to that model, from the same mislabelled table row as its MI300X sibling, and its note called the figure AMD's report.

8×B300 Cisco UCS C880A M8 · DeepSeek-R1

69,251.4 tok/s, counting generated tokens · MLPerf Inference v6.0 · offline

measured high confidence

8-GPU node; ~8,656 tok/s per GPU. Submission precision not independently confirmed.

8×B300 Cisco UCS C880A M8 · DeepSeek-R1

58,553.3 tok/s, counting generated tokens · MLPerf Inference v6.0 · server

measured high confidence

~7,319 tok/s per GPU; MLPerf latency-constrained server scenario.

GB300 NVL72 (full 72-GPU rack) · DeepSeek-R1

673,936 tok/s, counting generated tokens · MLPerf Inference v6.0 · offline

measured high confidence

Nebius submission; ~9,360 tok/s per GPU. CoreWeave submitted the same hardware class separately without publishing exact figures.

GB300 NVL72 (full 72-GPU rack) · DeepSeek-R1

575,580 tok/s, counting generated tokens · MLPerf Inference v6.0 · server

measured high confidence

~7,994 tok/s per GPU.

GB300 NVL72 (full 72-GPU rack) · GPT-OSS-120B

1,046,150 tok/s, counting generated tokens · MLPerf Inference v6.0 · offline

measured high confidence

~14,530 tok/s per GPU; identical figures in NVIDIA's and Nebius's recaps.

GB300 NVL72 (full 72-GPU rack) · GPT-OSS-120B

1,096,770 tok/s, counting generated tokens · MLPerf Inference v6.0 · server

measured high confidence

~15,233 tok/s per GPU.

GB200 NVL72, 64 of 72 GPUs (CoreWeave) · DeepSeek-R1

468,665 tok/s, counting generated tokens · MLPerf Inference v6.0 · offline

measured high confidence

Scale mismatch flagged: 16 nodes/64 GPUs, not a fully populated 72-GPU rack; ~7,323 tok/s per GPU.

8×MI355X · Llama 2 70B

103,480 tok/s, counting generated tokens · MLPerf Inference v6.0 (ROCm, FP4) · offline

measured high confidence

~12,935 tok/s per GPU; AMD states FP4 precision.

8×MI355X · Llama 2 70B

73,608 tok/s, counting generated tokens · MLPerf Inference v6.0 (ROCm, FP4) · interactive

measured high confidence

~9,201 tok/s per GPU; v6.0's Interactive scenario applies a tighter per-token latency SLA.

8×MI300X Supermicro · Llama 2 70B

27,804 tok/s, counting generated tokens · MLPerf Inference v5.1 · offline

measured high confidence

Neutral-ledger complement to the AMD-blog row (22,021 tok/s) — different harness and latency methodology; kept separate, never averaged.

8×MI300X Supermicro · Llama 2 70B

8,840 tok/s, counting generated tokens · MLPerf Inference v5.1 · interactive

measured high confidence

~1,105 tok/s per GPU.

8×MI325X QuantaGrid · Mixtral 8x7B

68,781 tok/s, counting generated tokens · MLPerf Inference v5.1 · offline

measured high confidence

A second 8×MI325X submission (ASUSTeK) reported 68,121 offline — close but distinct; both valid, not averaged.

8×B200 · Llama 3.1 405B

1,650 tok/s, counting generated tokens · MLPerf Inference v5.1 · offline

measured high confidence

~206 tok/s per GPU; audited non-preview complement to the v5.0 preview Llama-2-70B row on the same hardware class.

B200 (GPU count not disclosed on the compare view) · Qwen 3.5 397B-A17B

3,225.9 tok/s per accelerator, counting tokens this repository has not identified · InferenceX dashboard (FP8 default) · 69 tok/s/user interactivity point

estimated medium confidence

tok/s per GPU, interpolated between measured InferenceX runs (per the site's own methodology note); GPU count/TP not shown on the compare view. Which tokens it counts was not recorded: InferenceX derives three throughputs per GPU for every run, tput_per_gpu, input_tput_per_gpu and output_tput_per_gpu, and this row does not say which of them it was read from.

H200 · Qwen 3.5 397B-A17B

485.3 tok/s per accelerator, counting tokens this repository has not identified · InferenceX dashboard (FP8 default) · 69 tok/s/user interactivity point

estimated medium confidence

Same compare view as the B200 row; the ~6.6x gap is one interactivity slice, not a universal multiplier. Which of InferenceX's three per-GPU throughputs it is, total, input or output, was not recorded either.

2×GB300 (single node) · DeepSeek-V3.2-NVFP4

2,816 tok/s per accelerator, counting generated tokens · vLLM v0.14.1, CUDA 13.0 · mixed-context (ISL 2K/OSL 1K), TP2

measured medium confidence

tok/s per GPU on a 2-GPU dev node — deliberately NOT linked to the 72-GPU rack entity. Prefill-only on the same node: 7,360 tok/s per GPU. The post calls the 2,816 figure output throughput, so it counts generated tokens.

GB300 (TP8) · Kimi K3 (MXFP4)

111 tok/s per user while decoding · vLLM day-0 build, FP8 KV cache · single user (batch 1), 8K in / 1K out

vendor reported medium confidence

Per-user decode speed at batch 1, not aggregate serving throughput — the two differ by orders of magnitude and must not be compared directly. With DSpark speculative decoding the same configuration reaches 331 tokens/s per user, a 3.14x speed-up. Linked to the GB300 NVL72 with the 8 GPUs the run used. CORRECTED 2026-09-13: this note said the row was not linked because the catalog has no 8-GPU GB300 system. A run on part of a rack links the rack and records the GPUs it used, as the SGLang rows beside it do.

GB300 (TP16) · Kimi K3 (MXFP4)

118 tok/s per user while decoding · vLLM day-0 build, FP8 KV cache · single user (batch 1), 8K in / 1K out

vendor reported medium confidence

Per-user decode speed at batch 1 across 16 accelerators; 370 tokens/s with DSpark speculative decoding. Doubling tensor parallelism from TP8 buys only 6% single-user speed, which is the point: latency scales poorly with GPU count while aggregate throughput scales well. Linked to the GB300 NVL72 with the 16 GPUs the run used.

16× GB300: 1 PP8 prefill worker → 1 TP8 decode node (1P:1D) · Kimi K3

2,808 tok/s per accelerator, counting prompt and generated tokens · SGLang + Miles, prefill/decode disaggregated, fp4 arm · serving frontier at its throughput end, 8K in / 1K out, concurrency 1024

vendor reported medium confidence

The throughput end of the serving frontier in SGLang's Kimi K3 post: 2,808 tok/s per GPU at 18.7 tok/s per user, with a client concurrency limit of 1,024 printed beside the point. The figure counts prompt and generated tokens and divides by all 16 GPUs of the arm, the prefill worker's included, so the run served about 44,928 tok/s, about 4,992 of them generated; at 18.7 tok/s per user that is about 267 streams decoding at once, well inside the 1,024 limit. The post does not state that denominator; the post's source record sets out the measurements it is inferred from. The arm is tagged fp4, which the post never defines, and the figure says its points come from a separate sweep; the untagged PP8 → TP8 arm peaks at 2,715. That sweep's concurrency-1 point decodes 115.8 tok/s per user, beside the ~113 tok/s the post gives for batch 1 before speculation and far from the ~423 it gives with DSpark, so the sweep most likely ran without speculation; the post does not say. CORRECTED 2026-09-13: this row was sglang-k3-gb300-8gpu-batched and stored 22,464 tok/s, 2,808 decode tokens/s per GPU times 8 accelerators, left unlinked as an 8-GPU topology. The figure counts prompt tokens as well, and the run used 16 GPUs.

16× GB300: 1 PP8 prefill worker → 1 TP8 decode node (1P:1D) · Kimi K3

2,715 tok/s per accelerator, counting prompt and generated tokens · SGLang + Miles, prefill/decode disaggregated · serving frontier, untagged PP8 → TP8 arm at its peak, 8K in / 1K out

vendor reported medium confidence

The PP8 → TP8 · 1P:1D arm of the same frontier at its highest point, 2,715.1 tok/s per GPU at 18.74 tok/s per user, read from the SVG's data coordinates, whose axis ticks fit a straight line exactly; the figure prints no concurrency for it. It is the fp4 arm's topology without the fp4 tag. 16 GPUs × 2,715 = 43,440 tok/s counting prompt and generated tokens, about 4,827 of them generated, or about 258 streams at 18.74 tok/s. The per-GPU figure is divided by all 16 GPUs, the prefill worker's included: the post does not say so, and its source record infers it from the post's own measurements. The post names no precision and no speculation setting for its untagged arms. The caption calls arms within about 2% of one another indistinguishable.

32× GB300: 2 PP8 prefill workers → 2 DCP8 decode nodes (2P:2D) · Kimi K3

2,633 tok/s per accelerator, counting prompt and generated tokens · SGLang + Miles, prefill/decode disaggregated, TP8 + DCP8 decode · serving frontier, the DCP composition at its peak, 8K in / 1K out

vendor reported medium confidence

The caption's DCP composition, two PP8 prefill workers feeding two DCP8 decode nodes, at 2,633 tok/s per GPU; the figure plots it at 2,632.7 and 18.44 tok/s per user and prints no concurrency. Decode context parallelism shards the MLA cache inside each TP8 group, so a DCP8 decode node is still 8 GPUs and the arm is 32. 32 GPUs × 2,633 = 84,256 tok/s counting prompt and generated tokens, about 9,362 of them generated, or about 508 streams at 18.44 tok/s. The per-GPU figure is divided by all 32 GPUs, both prefill workers included: the post does not say so, and its source record infers it from the post's own measurements. The post names no precision and no speculation setting for its untagged arms. A single PP8 → DCP8 · 1P:1D arm peaks at 2,682.7 on 16 GPUs, within the 2% the caption calls indistinguishable.

8× GB300 (2×4 trays), TP8 + DCP8 + host-memory KV tier · Kimi K3

541 tok/s, counting tokens this repository has not identified · SGLang + Miles, DCP8, hicache ratio 2 · AgentX replay of coding-agent sessions, 48 concurrent sessions

vendor reported low confidence

The post's caption gives 541 tok/s at 48 sessions without saying what it counts, and only the alt text of the figure beside it calls that rate aggregate decode throughput. The row files it as unspecified, because the figure itself fits neither reading, as below. The run replays real coding-agent sessions on 2×4 GB300, 8 GPUs, with a host-memory KV tier at ratio 2 on both arms and decode context parallelism as the only difference; the TP8 arm collapses at 16 sessions, from 3,925 to 1,410 tok/s per GPU. The post names no precision and no speculation setting for the replay. The figure beside that text does not plot 541. Its axis counts prompt and generated tokens per GPU, and the DCP8 arm's 48-session point reads about 15,333 tok/s per GPU, about 122,660 across the 8 GPUs, at 8.9 tok/s per stream; 48 streams at 8.9 tok/s generate about 427 tok/s. 541 is neither that nor 122,660, and the post reconciles it with neither. The DCP8 curve peaks at 48 sessions and falls to about 12,750 at 56 and 8,520 at 64. Read from the SVG's data coordinates; the DCP8 panel has no axis labels of its own and is read on the TP8 panel's scale. CORRECTED 2026-09-13: this row said 541 was roughly 2.4% of the same hardware's short-context batched figure. That set a generated-token rate for 8 GPUs against 22,464, a prompt-inclusive per-GPU figure multiplied by 8, and the row linked no hardware although it ran on GB300.

8× B300 (1×8, prefill and decode unified) · Kimi K3 (MXFP4)

785 tok/s per accelerator, counting prompt and generated tokens · SGLang v0.5.18 @ 71de97b2, TP8 · cookbook Low-Latency cell, random 8,192 in / 1,024 out, concurrency 16

vendor reported medium confidence

Marked Verified in the cookbook: P50 TTFT 3,539 ms, P50 TPOT 19.47 ms and 785 tok/s per GPU at concurrency 16, with the native MXFP4 checkpoint, no --kv-cache-dtype flag, so the cache stays at the model's bf16, and no speculation. The per-GPU figure counts prompt and generated tokens: 16 streams at those medians generate 16 × 1,024 / (3.539 + 1,024 × 0.01947) = 697.9 tok/s, and 785 × 8 × 1,024 / 9,216 = 697.8. Measured with --random-range-ratio 1.0, --warmup-requests 64 and --flush-cache, so no request reuses another's prefix.

8× B300 (1×8, prefill and decode unified) · Kimi K3 (MXFP4)

1,395 tok/s per accelerator, counting prompt and generated tokens · SGLang v0.5.18 @ 71de97b2, TP8 + DCP8 · cookbook Balanced cell, random 8,192 in / 1,024 out, concurrency 64

vendor reported medium confidence

Marked Verified in the cookbook: P50 TTFT 11,635 ms, P50 TPOT 40.19 ms and 1,395 tok/s per GPU at concurrency 64, with the native MXFP4 checkpoint, decode context parallelism across the eight GPUs, a bf16 cache and no speculation. The per-GPU figure counts prompt and generated tokens: 64 × 1,024 / (11.635 + 1,024 × 0.04019) = 1,241.5 generated tok/s, against 1,395 × 8 × 1,024 / 9,216 = 1,240.0. The KDA state pool caps this recipe at 101 running requests, which is why the cookbook publishes no Balanced point past concurrency 64. Measured with --random-range-ratio 1.0, --warmup-requests 64 and --flush-cache.

8× B300 (1×8, prefill and decode unified) · Kimi K3 (MXFP4)

1,987 tok/s per accelerator, counting prompt and generated tokens · SGLang v0.5.18 @ 71de97b2, TP8 + DCP8, DSPARK speculative decoding · cookbook Balanced DSPARK cell, random 8,192 in / 1,024 out, concurrency 64

vendor reported low confidence

Marked Verified in the cookbook, with a simulated acceptance: P50 TTFT 12,038 ms, P50 TPOT 24.47 ms and 1,987 tok/s per GPU at concurrency 64, the Balanced recipe with DSPARK speculative decoding and --max-running-requests 256. The cookbook pins the draft acceptance length with SGLANG_SIMULATE_ACC_LEN=4.5, so the cell reports what DSPARK's block of 7 delivers at that acceptance, not an acceptance rate measured on this workload, and it asks readers to measure against the same recipe without speculation before adopting it. With DSPARK the KDA state pool caps admission at 68 running requests. The per-GPU figure counts prompt and generated tokens: 64 × 1,024 / (12.038 + 1,024 × 0.02447) = 1,766.7 generated tok/s, against 1,987 × 8 × 1,024 / 9,216 = 1,766.2.

8× MI350X (1×8, prefill and decode unified) · Kimi K3 (MXFP4)

462 tok/s per accelerator, counting prompt and generated tokens · SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV · cookbook Balanced cell, random 8,192 in / 1,024 out, concurrency 16

vendor reported low confidence

Marked Final Verification In Progress in the cookbook: P50 TTFT 6,331 ms, P50 TPOT 32.8 ms and 462 tok/s per GPU at concurrency 16. The MI350X and MI355X share one recipe: TP8 on ROCm with AITER's A8W4 MoE kernels, Triton attention, an fp8_e4m3 KV cache and, in this cell, no speculation. The per-GPU figure counts prompt and generated tokens: 16 × 1,024 / (6.331 + 1,024 × 0.0328) = 410.4 generated tok/s, against 462 × 8 × 1,024 / 9,216 = 410.7. Not linked to a catalog system, which carries the MI355X and not the MI350X.

8× MI350X (1×8, prefill and decode unified) · Kimi K3 (MXFP4)

813 tok/s per accelerator, counting prompt and generated tokens · SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV · cookbook Balanced cell, random 8,192 in / 1,024 out, concurrency 64

vendor reported low confidence

Marked Final Verification In Progress in the cookbook: P50 TTFT 18,665 ms, P50 TPOT 70.39 ms and 813 tok/s per GPU at concurrency 64, on the same recipe as the concurrency-16 cell. The per-GPU figure counts prompt and generated tokens: 64 × 1,024 / (18.665 + 1,024 × 0.07039) = 722.2 generated tok/s, against 813 × 8 × 1,024 / 9,216 = 722.7. Not linked to a catalog system, which carries the MI355X and not the MI350X.

8× MI350X (1×8, prefill and decode unified) · Kimi K3 (MXFP4)

913 tok/s per accelerator, counting prompt and generated tokens · SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV, DSPARK speculative decoding · cookbook Balanced DSPARK cell, random 8,192 in / 1,024 out, concurrency 64

vendor reported low confidence

Marked Final Verification In Progress in the cookbook: P50 TTFT 23,720 ms, P50 TPOT 31.45 ms and 913 tok/s per GPU at concurrency 64, with DSPARK speculative decoding on the shared MI350X/MI355X recipe. Here the medians do not reproduce the per-GPU figure: 64 × 1,024 / (23.720 + 1,024 × 0.03145) = 1,171.9 generated tok/s, against 913 × 8 × 1,024 / 9,216 = 811.6, a ratio of 0.69, while the same recipe's concurrency-16 DSPARK cell, 864 tok/s per GPU, agrees within 2%. The per-GPU figure is a whole-run total and the medians describe a typical request; the cookbook says neither which to trust nor whether this cell's acceptance length is simulated as the B300 cells' is. Not linked to a catalog system, which carries the MI355X and not the MI350X.

Source registry

NVIDIA H100 Tensor Core GPU

NVIDIA · accessed 2026-07-19 primary

Open source ↗
NVIDIA Supercharges Hopper: H200

NVIDIA Newsroom · accessed 2026-07-19 primary

Open source ↗
H200 price guide

Mercatus AI · accessed 2026-07-19 independent research

Open source ↗
Exclusive: prices of Nvidia B300 server

Reuters (Investing.com reprint) · accessed 2026-07-19 independent research

Open source ↗
AMD Instinct MI300X Platform Data Sheet

AMD · accessed 2026-07-19 primary

Open source ↗
AMD Instinct MI325X Platform Datasheet

AMD · accessed 2026-07-19 primary

Open source ↗
The Llama 3 Herd of Models

Meta AI · accessed 2026-08-21· published 2024-07-31 primary

Open source ↗
Llama model cards and license

Meta · accessed 2026-07-19 primary

Open source ↗
Mistral Large 3 model card

Mistral AI · accessed 2026-07-19 primary

Open source ↗
Mistral Large 3 675B Instruct repository

Mistral AI (Hugging Face) · accessed 2026-07-19 primary

Open source ↗
Claude API pricing

Anthropic · accessed 2026-08-28 primary

Open source ↗
Your data — data residency, supported endpoints and models

OpenAI · accessed 2026-09-10 primary

Open source ↗
GPT-5.6 Terra pricing

OpenAI · accessed 2026-09-13 primary

Open source ↗
Mistral API pricing

Mistral AI · accessed 2026-07-24 primary

Open source ↗
MLPerf Inference v5.0 results

MLCommons · accessed 2026-07-19 primary

Open source ↗
MLPerf Inference v4.1 results

MLCommons · accessed 2026-09-13 primary

Open source ↗
LLM inference performance

AMD ROCm · accessed 2026-09-13 primary

Open source ↗
MLPerf Inference v6.0 results

MLCommons · accessed 2026-07-24 primary

Open source ↗
MLPerf Inference v5.1 results

MLCommons · accessed 2026-07-24 primary

Open source ↗
MLPerf Inference LoadGen results.cc

MLCommons (LoadGen) · accessed 2026-09-13 primary

Open source ↗
AMD Instinct GPUs MLPerf Inference v6.0 submission

AMD ROCm · accessed 2026-07-24 primary

Open source ↗
MLPerf Inference v6.0 results on Blackwell and Blackwell Ultra

Nebius · accessed 2026-07-24 primary

Open source ↗
NVIDIA MLPerf Inference v6.0 recap

NVIDIA · accessed 2026-07-24 primary

Open source ↗
InferenceX (formerly InferenceMAX) open inference benchmark

SemiAnalysis · accessed 2026-07-24 independent research

Open source ↗
DeepSeek-V3.2 on GB300 (vLLM)

vLLM Project · accessed 2026-07-24 primary

Open source ↗
Claude pricing — tokenizer generation note

Anthropic · accessed 2026-08-29 primary

Open source ↗
Qwen3.5-397B-A17B model card

Qwen (Hugging Face) · accessed 2026-08-29 primary

Open source ↗
DeepSeek-V4-Pro model card

DeepSeek (Hugging Face) · accessed 2026-07-24 primary

Open source ↗
DeepSeek-V4-Flash model card

DeepSeek (Hugging Face) · accessed 2026-07-24 primary

Open source ↗
MiMo V2.5 open-source announcement

Xiaomi · accessed 2026-07-24 primary

Open source ↗
MiMo-V2.5-Pro model card

Xiaomi (Hugging Face) · accessed 2026-07-24 primary

Open source ↗
GLM-5.2 model card

Z.ai (Hugging Face) · accessed 2026-07-24 primary

Open source ↗
Kimi-K2.6 model card

Moonshot AI (Hugging Face) · accessed 2026-07-24 primary

Open source ↗
DeepSeek API pricing

DeepSeek · accessed 2026-09-13 primary

Open source ↗
MiMo API pricing (pay-as-you-go)

Xiaomi · accessed 2026-07-24 primary

Open source ↗
Alibaba Cloud Model Studio pricing

Alibaba Cloud · accessed 2026-07-24 primary

Open source ↗
Intel Gaudi 3 availability announcement

Intel · accessed 2026-07-24 primary

Open source ↗
Intel Gaudi 2/3 8x OAM UBB pricing

ServeTheHome · accessed 2026-07-24 independent research

Open source ↗
Introducing Claude Opus 5

Anthropic · accessed 2026-08-01· published 2026-07-24 primary

Open source ↗
Introducing Claude Sonnet 5

Anthropic · accessed 2026-08-01· published 2026-06-30 primary

Open source ↗
Advancing the price-performance frontier with GPT-5.6

OpenAI · accessed 2026-08-01· published 2026-07-30 primary

Open source ↗
OpenAI API changelog — GPT-5.6 Sol price reduction

OpenAI · accessed 2026-08-27· published 2026-08-21 primary

Open source ↗
moonshotai/Kimi-K3 model repository

Moonshot AI (Hugging Face) · accessed 2026-09-13· published 2026-07-27 primary

Open source ↗
deepseek-ai/DeepSeek-V4-Flash-0731 model repository

DeepSeek (Hugging Face) · accessed 2026-08-01· published 2026-07-31 primary

Open source ↗
DeepSeek API changelog

DeepSeek · accessed 2026-08-01 primary

Open source ↗
meta-llama/Llama-4-Maverick-17B-128E-Instruct

Meta (Hugging Face) · accessed 2026-08-01· published 2025-04-05 primary

Open source ↗
Llama 4 Community License Agreement

Meta · accessed 2026-08-01 primary

Open source ↗
Kimi K3 Is Here: Efficient Day-0 Support on vLLM

vLLM · accessed 2026-08-01· published 2026-07-27 primary

Open source ↗
SGLang and Miles Add Day-0 Support for Kimi K3

LMSYS Org · accessed 2026-09-13· published 2026-07-27 primary

Open source ↗
LGAI-EXAONE/K-EXAONE-2.0-750B-A37B

LG AI Research (Hugging Face) · accessed 2026-08-01· published 2026-07-31 primary

Open source ↗
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16

NVIDIA (Hugging Face) · accessed 2026-08-01· published 2026-06-04 primary

Open source ↗
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16

NVIDIA (Hugging Face) · accessed 2026-08-01· published 2026-03-11 primary

Open source ↗
MiniMaxAI/MiniMax-M3

MiniMax (Hugging Face) · accessed 2026-08-01· published 2026-06-07 primary

Open source ↗
tencent/Hy3

Tencent (Hugging Face) · accessed 2026-08-01· published 2026-07-06 primary

Open source ↗
CohereLabs/command-a-plus-05-2026-bf16

Cohere (Hugging Face) · accessed 2026-08-01 primary

Open source ↗
DeepSeek-V4-Pro GA Release

DeepSeek · accessed 2026-08-14· published 2026-08-13 primary

Open source ↗
deepseek-ai/DeepSeek-V4-Pro-0813 model repository

DeepSeek (Hugging Face) · accessed 2026-08-14· published 2026-08-13 primary

Open source ↗
Grok 4.6

xAI · accessed 2026-08-14· published 2026-08-12 primary

Open source ↗
meta-models/Muse-Glimmer-30B

Meta (Hugging Face) · accessed 2026-08-14· published 2026-08-10 primary

Open source ↗
Meta developer platform — pricing and rate limits

Meta · accessed 2026-08-14 primary

Open source ↗
Models — Meta Model API

Meta · accessed 2026-08-27 primary

Open source ↗
Qwen3.8-Max model documentation

Alibaba Cloud Model Studio · accessed 2026-09-13 primary

Open source ↗
Qwen/Qwen3.8-2.4T-A95B model repository

Qwen (Hugging Face) · accessed 2026-08-14 primary

Open source ↗
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

NVIDIA (Hugging Face) · accessed 2026-08-14· published 2026-08-11 primary

Open source ↗
Qwen/Qwen3.8-27B model repository

Qwen (Hugging Face) · accessed 2026-08-15· published 2026-08-13 primary

Open source ↗
Gemini 3.7 Flash model card

Google DeepMind · accessed 2026-08-29· published 2026-08-13 primary

Open source ↗
Artificial Analysis — independent LLM benchmarks

Artificial Analysis · accessed 2026-08-15 independent research

Open source ↗
Gemini API model deprecations

Google · accessed 2026-08-01 primary

Open source ↗
NVIDIA AI GPU pricing guide 2026

IntuitionLabs · accessed 2026-08-01 independent research

Open source ↗
DGX GB200 NVL72 User Guide — hardware

NVIDIA · accessed 2026-08-01 primary

Open source ↗
NVIDIA GB300 NVL72 by HPE — QuickSpecs

HPE · accessed 2026-08-01 independent research

Open source ↗
NVIDIA launches next-generation Rubin AI compute platform at CES 2026

ServeTheHome · accessed 2026-08-01· published 2026-01-05 independent research

Open source ↗
Vera Rubin: extreme co-design

SemiAnalysis · accessed 2026-08-01· published 2026-02-25 independent research

Open source ↗
AAI 2026: AMD launches Instinct MI400 Series GPUs

AMD Newsroom · accessed 2026-08-01· published 2026-07-23 primary

Open source ↗
AMD attacks the rack with Helios systems that rival Nvidia's

The Register · accessed 2026-08-01· published 2026-07-23 independent research

Open source ↗
Red Hat AI tops MLPerf Inference v6.0 with vLLM

Red Hat · accessed 2026-08-01 primary

Open source ↗
GLM-5.3: Frontier Coding with Emergent Cyber Capabilities

Z.ai · accessed 2026-08-15· published 2026-08-14 primary

Open source ↗
zai-org/GLM-5.2-FP8 model repository

Z.ai (Hugging Face) · accessed 2026-08-15 primary

Open source ↗
Qwen/Qwen3.8-27B-FP8 model repository

Qwen (Hugging Face) · accessed 2026-08-15 primary

Open source ↗
Qwen/Qwen3.8-2.4T-A95B-FP8 model repository

Qwen (Hugging Face) · accessed 2026-08-15 primary

Open source ↗
MiniMaxAI/MiniMax-M3-MXFP8 model repository

MiniMax (Hugging Face) · accessed 2026-08-15 primary

Open source ↗
mistralai/Mistral-Large-3-675B-Instruct-2512-BF16 model repository

Mistral AI (Hugging Face) · accessed 2026-08-15 primary

Open source ↗
mistralai/Mistral-Large-3-675B-Instruct-2512-NVFP4 model repository

Mistral AI (Hugging Face) · accessed 2026-08-15 primary

Open source ↗
Gemma 4 quantization-aware-training checkpoints

Google (Hugging Face) · accessed 2026-08-15 primary

Open source ↗
meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 model repository

Meta (Hugging Face) · accessed 2026-08-15 primary

Open source ↗
DeepSeek-V4 technical report, section 5.2.1 (FP4 quantization-aware training)

DeepSeek · accessed 2026-08-15 primary

Open source ↗
Prompt caching — OpenAI API documentation

OpenAI · accessed 2026-08-27 primary

Open source ↗
Prompt caching — Anthropic API documentation

Anthropic · accessed 2026-08-15 primary

Open source ↗
coding-usage — recorded agent session token accounting

Thomas Farfeleder · accessed 2026-08-15 independent research

Open source ↗
AtomicChat/Qwen3.8-27B-GGUF — quantizations with measured divergence

AtomicChat (Hugging Face) · accessed 2026-08-15 independent research

Open source ↗
nvidia/GLM-5.2-NVFP4

NVIDIA (Hugging Face) · accessed 2026-09-13 independent research

Open source ↗
nvidia/MiniMax-M3-NVFP4

NVIDIA (Hugging Face) · accessed 2026-09-13 independent research

Open source ↗
AtomicChat/DeepSeek-V4-Flash-0731-GGUF

AtomicChat (Hugging Face) · accessed 2026-08-15 independent research

Open source ↗
google/gemma-4-31b-it

Google (Hugging Face) · accessed 2026-08-27 primary

Open source ↗
google/gemma-4-e4b-it

Google (Hugging Face) · accessed 2026-08-27 primary

Open source ↗
zai-org/GLM-5.3-Flash

Z.ai (Hugging Face) · accessed 2026-08-25 primary

Open source ↗
zai-org/GLM-5.3-Flash-BF16

Z.ai (Hugging Face) · accessed 2026-08-27 primary

Open source ↗
GLM-5.3-Flash API documentation

Z.ai · accessed 2026-08-27 primary

Open source ↗
transformers modeling_glm5_next.py

Hugging Face · accessed 2026-08-27 primary

Open source ↗
Qwen/Qwen3.8-Flash-Next

Qwen (Hugging Face) · accessed 2026-08-27 primary

Open source ↗
Qwen/Qwen3.8-Flash-Next-FP8

Qwen (Hugging Face) · accessed 2026-08-27 primary

Open source ↗
Qwen Cloud model page: qwen3.8-flash

Alibaba · accessed 2026-08-27 primary

Open source ↗
transformers modeling_qwen4_exp.py

Hugging Face · accessed 2026-08-27 primary

Open source ↗
vLLM source at commit e52be1a6

vLLM · accessed 2026-09-13 primary

Open source ↗
SGLang source at commit fa663e72

SGLang · accessed 2026-09-13 primary

Open source ↗
vLLM recipe: zai-org/GLM-5.3-Flash

vLLM · accessed 2026-08-27 primary

Open source ↗
GLM-5.3-Flash Architecture Notes

Sebastian Raschka · accessed 2026-08-27· published 2026-08-26 independent research

Open source ↗
Muse Glimmer 30B Architecture Notes

Sebastian Raschka · accessed 2026-08-27· published 2026-08-11 independent research

Open source ↗
Qwen3.6-27B model repository

Qwen (Hugging Face) · accessed 2026-08-27 primary

Open source ↗
Qwen3.6-35B-A3B model repository

Qwen (Hugging Face) · accessed 2026-08-27 primary

Open source ↗
MiMo-V2.5 model repository

Xiaomi (Hugging Face) · accessed 2026-08-27 primary

Open source ↗
OpenAI model reference — gpt-5.6-cyber

OpenAI · accessed 2026-09-13 primary

Open source ↗
Claude context windows

Anthropic · accessed 2026-09-13 primary

Open source ↗
Claude models overview

Anthropic · accessed 2026-08-29 primary

Open source ↗
unsloth/Llama-4-Maverick-17B-128E-Instruct — config.json

Unsloth AI (Hugging Face) · accessed 2026-08-29 secondary

Open source ↗
RedHatAI/Llama-4-Maverick-17B-128E-Instruct-FP8 — config.json

Red Hat (Hugging Face) · accessed 2026-08-29 secondary

Open source ↗
nvidia/Qwen3.8-2.4T-A95B-NVFP4

NVIDIA (Hugging Face) · accessed 2026-08-29 independent research

Open source ↗
zai-org/GLM-5.3 model repository

Z.ai (Hugging Face) · accessed 2026-08-29 primary

Open source ↗
zai-org/GLM-5.3-BF16 model repository

Z.ai (Hugging Face) · accessed 2026-08-29 primary

Open source ↗
GLM-5.3 License

Z.ai (Hugging Face) · accessed 2026-08-29 primary

Open source ↗
Gemini 3.5 Flash-Lite model reference

Google · accessed 2026-09-13 primary

Open source ↗
Gemini 3.1 Flash-Lite model reference

Google · accessed 2026-09-13 primary

Open source ↗
Gemini 3.7 Flash model reference

Google · accessed 2026-09-13 primary

Open source ↗
tencent/Hy4-preview

Tencent (Hugging Face) · accessed 2026-08-29· published 2026-08-28 primary

Open source ↗
tencent/Hy4-preview-FP8

Tencent (Hugging Face) · accessed 2026-09-09· published 2026-08-27 primary

Open source ↗
ibm-granite/granite-4.2-30b

IBM (Hugging Face) · accessed 2026-08-29· published 2026-08-25 primary

Open source ↗
inclusionAI/Ling-3.0-flash

InclusionAI (Hugging Face) · accessed 2026-08-29· published 2026-08-02 primary

Open source ↗
stepfun-ai/Step-3.7-Flash

StepFun (Hugging Face) · accessed 2026-08-29· published 2026-05-23 primary

Open source ↗
mistralai/Mistral-Small-4-119B-2603

Mistral AI (Hugging Face) · accessed 2026-09-11· published 2026-03-16 primary

Open source ↗
google/gemma-4-E4B-it-qat-w4a16-ct

Google (Hugging Face) · accessed 2026-08-29· published 2026-06-04 primary

Open source ↗
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA · accessed 2026-09-08 primary

Open source ↗
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition

NVIDIA · accessed 2026-09-08 primary

Open source ↗
NVIDIA RTX PRO 6000 Blackwell Server Edition

NVIDIA · accessed 2026-09-08 primary

Open source ↗
NVIDIA RTX PRO 6000 Blackwell Server Edition Product Brief (SP-12355-001_v02)

NVIDIA · accessed 2026-09-09 primary

Open source ↗
ThinkSystem NVIDIA RTX PRO 6000 Blackwell Server Edition PCIe Gen5 GPU

Lenovo Press · accessed 2026-09-08 secondary

Open source ↗
Nvidia doubles RTX PRO 6000 Blackwell MSRP to a staggering $16,000

Tom's Hardware · accessed 2026-09-08 secondary

Open source ↗
GeForce RTX 5090 graphics card

NVIDIA · accessed 2026-09-08 primary

Open source ↗
NVIDIA RTX 6000 Ada Generation graphics card

NVIDIA · accessed 2026-09-08 primary

Open source ↗
AMD Radeon AI PRO R9700 GPU arrives October 27 at $1,299 for retail

TechPowerUp · accessed 2026-09-08 secondary

Open source ↗
Mac Studio technical specifications

Apple · accessed 2026-09-08 primary

Open source ↗
Apple introduces new Mac Studio with M5 Max and M5 Ultra

Apple · accessed 2026-09-08· published 2026-08-25 primary

Open source ↗
Claude Fable and Mythos 5.1

Anthropic · accessed 2026-09-08 primary

Open source ↗
Gemini 3.8 Flash and 3.8 Flash Cyber

Google · accessed 2026-09-08 primary

Open source ↗
Gemini Flash model card

Google DeepMind · accessed 2026-09-13 primary

Open source ↗
Qwen3.8-Max 0902

OpenRouter · accessed 2026-09-08 secondary

Open source ↗
Data residency — inference geo, pricing and model availability

Anthropic · accessed 2026-09-10 primary

Open source ↗
deepseek-ai/DeepSeek-V4.1-Flash model repository

DeepSeek (Hugging Face) · accessed 2026-09-11· published 2026-09-10 primary

Open source ↗
DeepSeek-V4.1 technical report

DeepSeek (Hugging Face) · accessed 2026-09-11· published 2026-09-10 primary

Open source ↗
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures

DeepSeek AI (ISCA '25) · accessed 2026-09-11· published 2025-05-14 independent research

Open source ↗
transformers — modular_gemma4.py and configuration_gemma4.py, the Gemma4ForConditionalGeneration reference implementation

Hugging Face · accessed 2026-09-11 primary

Open source ↗
DeepSeek-V4.1-Flash release

DeepSeek · accessed 2026-09-13· published 2026-09-10 primary

Open source ↗
Gemini 3.8 Flash model reference

Google · accessed 2026-09-13 primary

Open source ↗
Gemini API token guide

Google · accessed 2026-09-13 primary

Open source ↗
Z.ai API reference — Chat Completion

Z.ai · accessed 2026-09-13 primary

Open source ↗
Claude thinking documentation

Anthropic · accessed 2026-09-13 primary

Open source ↗
Claude batch processing documentation

Anthropic · accessed 2026-09-13 primary

Open source ↗
Claude fast mode documentation

Anthropic · accessed 2026-09-13 primary

Open source ↗
OpenAI token counting guide

OpenAI · accessed 2026-09-13 primary

Open source ↗
Gemini 2.5 Pro model reference

Google · accessed 2026-09-13 primary

Open source ↗
Gemini 3.1 Pro Preview model reference

Google · accessed 2026-09-13 primary

Open source ↗
Gemini 3.6 Flash model reference

Google · accessed 2026-09-13 primary

Open source ↗
Gemini API thinking guide

Google · accessed 2026-09-13 primary

Open source ↗
Kimi K3 quickstart

Moonshot AI · accessed 2026-09-13 primary

Open source ↗
Grok 4.6 developer guide

xAI · accessed 2026-09-13 primary

Open source ↗
Qwen3.7-Max model documentation

Alibaba Cloud Model Studio · accessed 2026-09-13 primary

Open source ↗
Qwen3.7-Plus model documentation

Alibaba Cloud Model Studio · accessed 2026-09-13 primary

Open source ↗
Qwen3.8-Flash model documentation

Alibaba Cloud Model Studio · accessed 2026-09-13 primary

Open source ↗
DeepSeek API reference — Create Chat Completion

DeepSeek · accessed 2026-09-13 primary

Open source ↗
MiMo API reference — OpenAI-compatible chat completions

Xiaomi · accessed 2026-09-13 primary

Open source ↗
MiMo-V2.5-Pro model page

Xiaomi · accessed 2026-09-13 primary

Open source ↗
Mistral Medium 3.5 model page

Mistral AI · accessed 2026-09-13 primary

Open source ↗
Mistral API reference — Chat Completion

Mistral AI · accessed 2026-09-13 primary

Open source ↗

Where every figure comes from

One row per claim in the catalog: what it is about, how strong the evidence is, which sources back it, and — where the figure was derived rather than read off a page — the derivation itself, so you can redo it and disagree. The five grades are defined under Methodology.

338 of 338 claims · 68 are derived or assumed

  • H100 SXM Specification and price basis official claim high confidence

    Official specifications.

  • H200 8-GPU system Specification and price basis assumption medium confidence

    Memory/bandwidth use official H200 values; the 8-GPU price range is independent research and power is a planning assumption.

  • DGX B200 Specification and price basis official claim high confidence

    Official node specifications; price intentionally missing.

  • DGX B300 Specification and price basis estimated medium confidence

    CORRECTED at the 2026-08-01 refresh. This entry previously carried $550,000, a figure that re-verification could not trace to any source; current independent reporting puts an 8-GPU B300 node at $300k-350k baseline and up to $400-500k from resellers, so the point value is now $350k with the spread recorded. Memory (2,304 GB), 64 TB/s aggregate HBM bandwidth and 14.5 kW are all exact matches to NVIDIA's own DGX B300 User Guide. Note that the 14.4 TB/s figure circulating in secondary write-ups is the NVLink switch fabric, not memory bandwidth. Deployment multiplier is an editable planning assumption.

  • GB200 NVL72 Specification and price basis official claim medium confidence

    CORRECTED at the 2026-08-01 refresh: power is now NVIDIA's own approximately 120 kW from the DGX GB200 NVL72 User Guide, replacing the 132 kW previously recorded. CONFLICT preserved: infrastructure partners co-designing reference racks with NVIDIA (Vertiv, Schneider Electric, HPE) provision for up to 132 kW nominal with far higher transient peaks, so 132 kW is the right number for facility sizing and 120 kW for the rack itself. Memory and fabric are official; the 13.4 TB rack aggregate is NVIDIA's own published figure and deliberately is not 72 x a per-GPU value.

  • GB300 NVL72 Specification and price basis estimated low confidence

    CORRECTED at the 2026-08-01 refresh. The previous $3.0M simply reused the GB200 estimate, while reporting consistently puts GB300 at a premium over its predecessor ($3.0-4.0M, with one analyst note derived from a large customer order suggesting $3.7-4.0M); the midpoint is now recorded with the spread. Power moved from 132 kW to 135 kW: NVIDIA publishes no rack figure for GB300 at all, OEM reference documentation from HPE and Lenovo converges on 132-135 kW nominal, and other trackers report 140-142 kW — treat 132-142 kW as the honest range. Memory is derived as 72 x 288 GB; NVIDIA's own page rounds it to 20 TB.

  • Vera Rubin NVL72 (VR200) Specification and price basis estimated low confidence

    NAMING: NVIDIA announced this rack as "NVL144" at GTC 2025 counting compute dies, then as "NVL72" at CES 2026 counting GPU packages. They are the same rack, not two products. Memory (288 GB HBM4 x 72), bandwidth and NVLink 6 fabric are vendor figures. The 190 kW power draw is NOT vendor-published — it is the low end of a supply-chain analyst range of roughly 190 kW (Max-Q) to 230 kW (Max-P), so treat facility power here as an editable planning assumption. Declared in full production at CES 2026 with volume availability in the second half of 2026. DOWNGRADED 2026-08-29: this record was graded official-claim at medium confidence while citing no primary source at all — both sourceIds are independent-research, ServeTheHome and SemiAnalysis, so nothing here was read from NVIDIA. The gb300-nvl72 record carries the identical caveat about an unpublished rack power figure, cites one primary source more than this one, and is graded estimated at low confidence; the entry with the weaker evidence held the stronger grade. Nothing about the figures changed — 288 GB HBM4 x 72 and the NVLink 6 fabric are still what NVIDIA announced at CES 2026 and what those two outlets reported — only the claim this catalog makes about how it knows them.

  • Instinct MI455X Helios rack Specification and price basis official claim medium confidence

    NAMING: AMD's own press releases call this generation both "MI450 Series" (the OpenAI and Anthropic partnership announcements) and "MI400 Series" with the MI455X as flagship (the July 2026 launch). Same hardware. 432 GB HBM4 per accelerator across 72 accelerators is a vendor figure, as is the 260 TB/s UALink-over-Ethernet fabric. CONFLICT preserved on power: coverage of the same launch event cites both roughly 140 kW and 225-245 kW of bus-bar capacity under load; AMD has published no single authoritative rack figure, so 190 kW is recorded as a midpoint planning assumption and should be edited to whatever a vendor quote states. Declared in full production 2026-07-23 with shipments from Q3 2026.

  • Instinct MI300X 8-GPU platform Specification and price basis assumption medium confidence

    Memory and bandwidth are official. The data sheet's GPU-only 6 kW is converted to an editable 8.5 kW full-node planning value for TCO.

  • Instinct MI325X 8-GPU platform Specification and price basis assumption medium confidence

    Memory and bandwidth are official. The data sheet's GPU-only 8 kW is converted to an editable 10.5 kW full-node planning value for TCO.

  • Instinct MI355X 8-GPU platform Specification and price basis assumption medium confidence

    Official GPU specifications; node price/power are editable planning assumptions.

  • Gaudi 3 8-OAM node Specification and price basis assumption medium confidence

    128 GB HBM2e and 900 W per accelerator are official; the 9.5 kW full-node figure converts the 7.2 kW accelerator-only TDP into an editable planning value (same convention as the AMD nodes).

  • RTX PRO 6000 Blackwell Specification and price basis official claim high confidence

    96 GB GDDR7 ECC, 1,792 GB/s and 600 W are NVIDIA's own specifications; the acquisition figure is not, and is recorded as a range because no vendor price exists to check it against. Roughly a doubling since launch, attributed by the same reporting to a GDDR7 shortage. Carries no NVLink: the datasheet lists PCIe 5.0 x16 as the whole of its system interface, and the Server Edition's product brief states 'NVIDIA NVLink: Not supported' outright for the same silicon. NVLink left this product line two generations ago — the RTX A6000 had a bridge for two cards at 112 GB/s, the RTX 6000 Ada that replaced it did not — so a second card in the same box talks over the host bus at 128 GB/s, not over a fabric.

  • RTX PRO 6000 Blackwell Max-Q Specification and price basis official claim high confidence

    Same 96 GB and the same 1,792 GB/s as the full-power Workstation Edition at half the board power: a power-capped bin for slim chassis, not a cut-down memory system. Compute throughput at 300 W is lower, and this catalog holds no measurement of by how much. Carries no NVLink: the datasheet lists PCIe 5.0 x16 as the whole of its system interface, and the Server Edition's product brief states 'NVIDIA NVLink: Not supported' outright for the same silicon. NVLink left this product line two generations ago — the RTX A6000 had a bridge for two cards at 112 GB/s, the RTX 6000 Ada that replaced it did not — so a second card in the same box talks over the host bus at 128 GB/s, not over a fabric.

  • 4× RTX PRO 6000 Blackwell Max-Q Specification and price basis assumption medium confidence

    Four is the count NVIDIA itself names: the Max-Q datasheet says "Scale up to four RTX PRO 6000 Max-Q GPUs", and the 600 W Workstation Edition datasheet makes no such claim, which is why this build is the 300 W part. Memory and bandwidth are four times the card’s own published figures. Power is board power for the four cards and excludes the host, matching how the single-card entries are recorded. The H100 page is cited for one thing only, and a review was right that nothing said so: NVIDIA prints "NVLink: 900GB/s | PCIe Gen5: 128GB/s" in one spec table, which fixes both the figure and the convention it is on, since the 900 is the bidirectional total this catalog already stores for H100 SXM. The A6000 page dates the loss of NVLink from the line.

  • RTX PRO 6000 Blackwell Server Edition Specification and price basis official claim high confidence

    The rack card carries the same 96 GB GDDR7 as the two workstation editions but publishes 1,597 GB/s against their 1,792 GB/s — a lower memory clock, stated on both NVIDIA's own page and an OEM product guide, so it is not a typo on one page. Decode is bound by that figure, which makes the server part the slowest of the three despite the datacenter label. Carries no NVLink: the datasheet lists PCIe 5.0 x16 as the whole of its system interface, and the Server Edition's product brief states 'NVIDIA NVLink: Not supported' outright for the same silicon. NVLink left this product line two generations ago — the RTX A6000 had a bridge for two cards at 112 GB/s, the RTX 6000 Ada that replaced it did not — so a second card in the same box talks over the host bus at 128 GB/s, not over a fabric.

  • RTX 6000 Ada Generation Specification and price basis official claim medium confidence

    Previous generation and still sold new. It is here because "the 96 GB one or the 48 GB one" describes two generations rather than two configurations of one card: 96 GB is Blackwell, 48 GB is Ada, and the Ada part also has barely half the memory bandwidth.

    NVIDIA ↗ read 2026-09-08
  • GeForce RTX 5090 Specification and price basis official claim medium confidence

    Same GB202 die and the same 1,792 GB/s as the RTX PRO 6000, with a third of the memory. That pairing is why it is listed: on a consumer card it is capacity that runs out first, not bandwidth. NVIDIA's own page rendered neither price nor bandwidth in this session's fetch, so both are corroborated rather than read off the vendor.

    NVIDIA ↗ read 2026-09-08
  • DGX Spark Specification and price basis official claim high confidence

    128 GB of unified LPDDR5x buys capacity no discrete desktop card offers, at 273 GB/s — about a seventh of an RTX PRO 6000. Decode is memory-bound, so this box fits models it cannot serve quickly, and the catalog holds no throughput measurement for it. Power is the 240 W supply rating; the GB10's own TDP is 140 W.

    NVIDIA ↗ read 2026-09-08
  • Radeon AI PRO R9700 Specification and price basis official claim high confidence

    32 GB GDDR6 at 640 GB/s and 300 W, from AMD's own product page. The price is press reporting of the launch figure, not a line on that page.

    AMD ↗TechPowerUp ↗ read 2026-09-08
  • Mac Studio (M5 Max, 128 GB) Specification and price basis official claim medium confidence

    Apple publishes two M5 Max bandwidths, 460 GB/s for the 32-core GPU and 614 GB/s for the 40-core; this entry is the higher bin, so a cheaper Mac Studio is also a slower one. Every throughput profile in this catalog was measured on a CUDA or ROCm serving stack and none of them transfers to Metal, so this system carries no curated rate and needs a manual one.

    Apple ↗Apple ↗ read 2026-09-08
  • Kimi K3 Parameters, licence and weights official claim high confidence

    Parameter counts, context, quantization scheme and licence read from the published model card, config.json and LICENSE. The 1,560.86 GB checkpoint size is the safetensors index's own metadata.total_size, 1,560,860,324,864 bytes, rather than a figure quoted from a secondary write-up. Moonshot's technical blog, not the model card, carries the deployment floor: "we recommend deploying Kimi K3 on supernode configurations with 64 or more accelerators." CORRECTED 2026-09-13: this note gave the size as 1,560.94 GB, recomputed from the repository's shard byte sizes, after the model's own notes had moved to the index.

  • Kimi K3 KV-cache geometry estimated high confidence

    Derived from config.json: 24 Gated-MLA layers cache kv_lora_rank 512 + qk_rope_head_dim 64 = 576 elements per token, and 69 Kimi Delta Attention layers hold a fixed state of linear_attn_config num_heads 96 x head_dim 128 x 128 = 1,572,864 elements at fp32, plus three short-convolution windows of 96 x 128 x (short_conv_kernel_size 4 - 1) = 36,864 elements each at bf16: 6,512,640 bytes per layer, 449.37 MB per sequence. VALIDATED against SGLang's published figures: 24 x 576 x 2 = 27,648 bytes per token at bf16, or 27.0 KiB, against their stated "MLA KV block per token: 27 KB"; and one rank of eight holds 69 x (12 x 128 x 128 x 4 + 3 x 12 x 128 x 3 x 2) = 56,171,520 bytes of recurrent state against their stated "KDA state per request: 54 MB under TP=8". That is the sum SGLang itself computes — its ratio calculator for this model multiplies out exactly these terms — and it rounds to 54 in binary megabytes, 53.57 MiB, where decimal megabytes would read 56.17. Note that mla_use_nope does not remove the rope channel from the cache: the rotation is skipped, the 64 dimensions are still stored. The MLA latent is a single KV head, so plain tensor parallelism replicates it on every rank; decode context parallelism shards it, which is why SGLang's 48-concurrent-session result used DCP8. CORRECTED 2026-08-21: this profile defaulted to a one-byte element, alone among the twenty-five profiles here, on the reading that SGLang serves K3 with an fp8 cache. It does not. SGLang's own published block size — one MLA KV block per token of about 27 KB across all 24 MLA layers — is the bf16 figure, since 24 x 576 x 2 = 27,648 bytes and one byte would give 13.5 KiB. Their cookbook offers --kv-cache-dtype fp8_e4m3 as a lever that halves KV bytes per token, which is only meaningful from a two-byte baseline, and the checkpoint's own quantization_config carries kv_cache_scheme: null, so nothing in the release prescribes fp8. The default model on the front page was therefore showing half its cache: 3.24 GB where it should read 6.00 GB, and 133 concurrent sequences where 71 fit. The fp8 path is real and stays available as a control; it is not the default. CORRECTED 2026-09-13: the state was 69 x 1,720,320 elements at fp32, 474.81 MB per sequence. That counted each short-convolution window as the whole kernel of four inputs and stored it at fp32 beside the state, but a window keeps short_conv_kernel_size - 1 = 3 inputs, and both runtimes hold it at the activation width while the state stays at float32. vLLM e52be1a6 sizes the window conv_kernel_size - 1 + num_spec in kda_state_shape (model_executor/layers/mamba/mamba_utils.py:298) and takes it at the model dtype in kda_state_dtype (:133); SGLang fa663e72 sizes it conv_kernel_size - 1 in KimiLinearStateShape.create (srt/configs/mamba_utils.py:324) at SGLANG_MAMBA_CONV_DTYPE, default bfloat16 (srt/environ.py:1348). The validation sentence used to put 69 x 12 x 128 x 128 x 4 = 54.26 MB against the same published 54 and conclude that "their figure covers the recurrent state alone, so the convolutions here add about 9% on top of it". It matched for the wrong reason: 54.26 is the fp32 state alone in decimal megabytes, SGLang counts window and state together (mamba_cache_per_req, srt/configs/mamba_utils.py:116), and at their true size the windows add 3.52% to the state. The split by head is exact while the tensor-parallel size divides the 96 heads: 1, 4 and 8 do and 72 does not, and both runtimes' state shapes refuse a size that does not. vLLM's --mamba-ssm-cache-dtype bfloat16 and SGLang's --mamba-ssm-dtype bfloat16 each halve the state.

  • Kimi K3 Attention mechanism official claim high confidence

    config.json linear_attn_config: 69 KDA layers and 24 full Gated-MLA layers, 93 total.

  • Kimi K3 Expert routing official claim high confidence

    config.json: num_experts 896, num_experts_per_token 16, num_shared_experts 2, moe_intermediate_size 3072.

  • Kimi K3 Weights availability official claim high confidence
  • Llama 3.1 405B Parameters, licence and weights official claim high confidence
  • Llama 3.1 405B KV-cache geometry estimated high confidence

    2 x num_key_value_heads 8 x head_dim 128 = 2,048 elements per token per layer across 126 layers: 516.10 MB per 1,000 tokens at bf16. VALIDATED against a published figure — Table 1 of DeepSeek's ISCA 2025 hardware paper tabulates Llama-3.1-405B at 516.096 KB per token at bf16, which this reproduces exactly. Kept in the catalog after retirement precisely because it is the baseline every newer architecture is measured against: the cheapest model here costs a hundredth of this per token of context.

  • Llama 3.1 405B Attention mechanism official claim high confidence

    126 layers, 8 KV heads, head dimension 128, no sliding window and no compression. Read from Table 3 of Meta's own Llama 3 paper rather than from config.json: Meta gates the model repository, so the config file cannot be cited as something a reader can check. The previously cited source here was the llama-models GitHub repository, which carries licence and policy text and no architecture table at all — the numbers were right and the citation did not support them. Five independently uploaded derivatives of the same checkpoint agree with the paper on every field.

  • Llama 3.1 405B Weights availability official claim high confidence
  • Llama 4 Maverick Parameters, licence and weights official claim high confidence
  • Llama 4 Maverick KV-cache geometry estimated high confidence

    2 x num_key_value_heads 8 x head_dim 128 = 2,048 elements per token per layer. config.json no_rope_layers marks every fourth layer as global NoPE — 12 of 48 — and attention_chunk_size 8192 caps the other 36. Past 8,192 tokens the marginal cost is 49.15 MB per 1,000 tokens against 196.61 MB inside the chunk, but the chunked layers saturate at 1.21 GB per sequence, the largest fixed term in this catalog.

  • Llama 4 Maverick Attention mechanism official claim medium confidence

    config.json: 40 query heads over 8 KV heads at head dimension 128, across 48 layers. no_rope_layers marks 12 of those 48 as global NoPE full-attention and the remaining 36 as chunked at attention_chunk_size 8,192; transformers reads a 1 in that array as a RoPE chunked layer and a 0 as a NoPE global one. CORRECTED 2026-08-21: this note previously said the chunk size and the chunked-layer set could not be read because Meta gates the repository. That was true of the gated file and false of the entry — the kvCache block beside it had already sized the cache from an ungated mirror two weeks earlier, so one half of this model's record contradicted the other. Re-verified against two independently uploaded mirrors of the same checkpoint, which agree on every field.

  • Llama 4 Maverick Expert routing official claim high confidence

    config.json: num_local_experts 128, num_experts_per_tok 1, one shared expert per MoE layer. interleave_moe_layer_step 2, so only 24 of the 48 layers are MoE at all.

  • Llama 4 Maverick Weights availability official claim high confidence
  • Llama 4 Maverick Precision: FP8 (Meta) official claim high confidence

    84 shards summing to 416,755,504,016 bytes.

  • Qwen3.8-2.4T-A95B Parameters, licence and weights official claim medium confidence
  • Qwen3.8-2.4T-A95B KV-cache geometry estimated high confidence

    Model card: "23 x (3 x (Gated DeltaNet -> MoE) -> 1 x (Gated Attention -> MoE))" over 92 layers. The 23 growing layers cache 2 x num_key_value_heads 4 x head_dim 256 = 2,048 elements per token — 94.21 MB per 1,000 tokens. The 69 linear layers hold linear_num_value_heads 128 x linear_value_head_dim 128 x linear_key_head_dim 128 = 2,097,152 state elements at fp32 plus a convolution window over q and k at linear_num_key_heads 16 and v at the 128 value heads, (2 x 16 x 128 + 128 x 128) x (linear_conv_kernel_dim 4 - 1) = 61,440 elements at bf16, about 587 MB per sequence: the largest constant term in this catalog, and a real constraint on how many sequences a deployment can hold. The recurrent-state layout is the one both runtimes allocate for Gated DeltaNet: vLLM e52be1a6's gated_delta_net_state_shape (model_executor/layers/mamba/mamba_utils.py:274) and SGLang fa663e72's Mamba2StateShape.create with the key heads as its groups (srt/configs/qwen3_next.py:300). CORRECTED 2026-09-13: the convolution state was 81,920 elements at fp32 and the constant 601.42 MB per sequence. The window keeps linear_conv_kernel_dim - 1 = 3 inputs, not 4, and both runtimes store it at the model's bfloat16 while the state follows mamba_ssm_dtype, float32: vLLM takes the window at the model dtype and the state from mamba_ssm_cache_dtype (_mamba_state_dtype, model_executor/layers/mamba/mamba_utils.py:98), which model_executor/models/config.py sets from this checkpoint's mamba_ssm_dtype; SGLang takes the window at SGLANG_MAMBA_CONV_DTYPE, default bfloat16 (srt/environ.py:1348), and the state from the same field. The split by head is exact while the tensor-parallel size divides both the 16 key heads and the 128 value heads: 1, 4 and 8 do and 72 does not. vLLM's --mamba-ssm-cache-dtype bfloat16 and SGLang's --mamba-ssm-dtype bfloat16 each halve the state. UPGRADED 2026-09-13: confidence medium to high. The layout was inferred from the Gated DeltaNet formulation; it is now read from the code both runtimes allocate it with, and every field it uses from this checkpoint's config.json.

  • Qwen3.8-2.4T-A95B Attention mechanism official claim high confidence

    config.json layer_types over 92 layers: 69 linear_attention, 23 full_attention, full_attention_interval 4.

  • Qwen3.8-2.4T-A95B Expert routing official claim high confidence

    config.json uses non-standard key names: num_experts 512 (not n_routed_experts), num_experts_per_tok 10, and the shared expert appears only as shared_expert_intermediate_size 2048.

  • Qwen3.8-2.4T-A95B Weights availability official claim high confidence
  • Qwen3.8-2.4T-A95B Precision: FP8 (Alibaba) official claim high confidence

    safetensors index metadata.total_size 2,496,066,252,544 bytes.

  • Qwen3.8-2.4T-A95B Precision: NVFP4 (NVIDIA) official claim high confidence

    safetensors index metadata.total_size 1,444,420,107,432 bytes across 569,081 tensors, read through /resolve/ because /raw/ returns a git-LFS pointer for this repository. Graded an official claim rather than a measurement: NVIDIA publishes the artefact and its size, not an accuracy evaluation against the bf16 baseline, so nothing here says what the quantization cost in quality.

  • Qwen3.8-27B Parameters, licence and weights official claim high confidence
  • Qwen3.8-27B KV-cache geometry estimated high confidence

    Only 16 layers grow with context: 2 x num_key_value_heads 4 x head_dim 256 = 2,048 elements per token per layer, or 64 KiB per token, or 65.54 MB per 1,000 tokens across the model at bf16 — a twentieth of what a same-sized model with attention on every layer would hold. The 48 linear layers hold a fixed state: linear_num_value_heads 48 x linear_key_head_dim 128 x linear_value_head_dim 128 = 786,432 elements at fp32, plus a convolution window over q and k at linear_num_key_heads 16 and v at the 48 value heads, (2 x 16 x 128 + 48 x 128) x (linear_conv_kernel_dim 4 - 1) = 30,720 elements at bf16: 3,207,168 bytes per layer, 153.94 MB per sequence. The per-token slope is read straight from the config. The recurrent-state layout is the one both runtimes allocate for Gated DeltaNet — vLLM e52be1a6's gated_delta_net_state_shape (model_executor/layers/mamba/mamba_utils.py:274) and SGLang fa663e72's Mamba2StateShape.create with the key heads as its groups (srt/configs/qwen3_next.py:300) — and no vendor number exists to validate it against, unlike Kimi K3's. An error there shifts the constant term and not the slope. CORRECTED 2026-09-13: the convolution term was 40,960 elements at fp32, 827,392 elements per layer and 158.86 MB per sequence. The window keeps linear_conv_kernel_dim - 1 = 3 inputs, not 4, and both runtimes store it at the model's bfloat16 while the state follows mamba_ssm_dtype, float32: vLLM takes the window at the model dtype and the state from mamba_ssm_cache_dtype (_mamba_state_dtype, model_executor/layers/mamba/mamba_utils.py:98), which model_executor/models/config.py sets from this checkpoint's mamba_ssm_dtype; SGLang takes the window at SGLANG_MAMBA_CONV_DTYPE, default bfloat16 (srt/environ.py:1348), and the state from the same field. The split by head is exact while the tensor-parallel size divides both the 16 key heads and the 48 value heads: 1, 4 and 8 do and 72 does not. vLLM's --mamba-ssm-cache-dtype bfloat16 and SGLang's --mamba-ssm-dtype bfloat16 each halve the state. UPGRADED 2026-09-13: confidence medium to high. The layout was inferred from the Gated DeltaNet formulation; it is now read from the code both runtimes allocate it with, and every field it uses from this checkpoint's config.json.

  • Qwen3.8-27B Attention mechanism official claim high confidence

    config.json layer_types: 48 linear_attention and 16 full_attention, full_attention_interval 4. The model card describes it as "16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))".

  • Qwen3.8-27B Weights availability official claim high confidence
  • Qwen3.8-27B Precision: FP8 (Alibaba) official claim medium confidence

    Summed from the repository file listing: 30,866,866,928 bytes across 64 layer shards plus mtp.safetensors and outside.safetensors. No metadata.total_size is published for this repository.

  • Qwen3.8-27B Precision: Q6_K (AtomicChat) measured high confidence

    Size read from the repository file listing. The accuracy figure is the quantizer’s own published measurement, reproducible from the logs, calibration corpus and reference logits they publish alongside it.

  • Qwen3.8-27B Precision: Q4_K_M (AtomicChat) measured high confidence

    Size read from the repository file listing. The accuracy figure is the quantizer’s own published measurement, reproducible from the logs, calibration corpus and reference logits they publish alongside it.

  • Qwen3.8-27B Precision: IQ3_S (AtomicChat) measured high confidence

    Size read from the repository file listing. The accuracy figure is the quantizer’s own published measurement, reproducible from the logs, calibration corpus and reference logits they publish alongside it.

  • Muse Glimmer 30B Parameters, licence and weights official claim high confidence
  • Muse Glimmer 30B KV-cache geometry estimated high confidence

    2 x num_key_value_heads 2 x head_dim 128 = 512 elements per token per layer. Two KV heads is aggressive even by grouped-query standards, and it is the main reason this model costs 13.31 MB per 1,000 tokens at long context — a twenty-fifth of Hunyuan Hy3. The cache stops dividing beyond two accelerators. CROSS-CHECKED 2026-08-27 against Sebastian Raschka's independently published BF16 cache-per-token figures: counting every layer as growing, before the window cap, this geometry gives 52.00 KiB per token against his 52 KiB. An exact match.

  • Muse Glimmer 30B Attention mechanism official claim high confidence

    Model card: "Attention pattern | [Local, Local, Local, Global] repeating"; "RoPE (theta = 500,000), local layers only". config.json layer_types puts global attention at every fourth layer — 13 global and 39 sliding of 52 — with sliding_window 2048 and only 2 KV heads.

  • Muse Glimmer 30B Weights availability official claim high confidence
  • Nemotron 3.5 Lightning 30B-A3B Parameters, licence and weights official claim high confidence
  • Nemotron 3.5 Lightning 30B-A3B KV-cache geometry estimated high confidence

    6 attention layers at 512 elements per token: 6.14 MB per 1,000 tokens. State-space layers hold mamba_num_heads 64 x mamba_head_dim 64 x ssm_state_size 128 = 524,288 elements at fp32 plus a convolution window of (64 x 64 + 2 x n_groups 8 x 128) x (conv_kernel 4 - 1) = 18,432 elements at bf16, about 49 MB per sequence. That is the layout both runtimes allocate for Mamba-2: vLLM e52be1a6's mamba2_state_shape (model_executor/layers/mamba/mamba_utils.py:195) and SGLang fa663e72's Mamba2StateShape.create (srt/configs/mamba_utils.py:195). CORRECTED 2026-09-13: the convolution state was 24,576 elements at fp32, about 50 MB per sequence. A Mamba-2 window keeps conv_kernel - 1 = 3 inputs, and both runtimes store it at the model's bfloat16 while the state stays at float32: vLLM takes the window at the model dtype and the state from mamba_ssm_cache_dtype (_mamba_state_dtype, model_executor/layers/mamba/mamba_utils.py:98), which model_executor/models/config.py sets from this checkpoint's own mamba_ssm_cache_dtype; SGLang takes the window at SGLANG_MAMBA_CONV_DTYPE, default bfloat16 (srt/environ.py:1348), and the state at its float32 default. The split by head is exact while the tensor-parallel size divides both the 8 groups and the 64 heads: 1, 4 and 8 do and 72 does not, and vLLM's MambaMixer2 asserts both (model_executor/layers/mamba/mamba_mixer2.py:305 and :309, the second waived only for a single group). vLLM's --mamba-ssm-cache-dtype bfloat16 and SGLang's --mamba-ssm-dtype bfloat16 each halve the state.

  • Nemotron 3.5 Lightning 30B-A3B Attention mechanism official claim high confidence

    config.json layers_block_type over 52 entries: 23 mamba, 23 moe and 6 attention.

  • Nemotron 3.5 Lightning 30B-A3B Weights availability official claim high confidence
  • K-EXAONE 2.0 750B-A37B Parameters, licence and weights official claim high confidence
  • K-EXAONE 2.0 750B-A37B KV-cache geometry estimated high confidence

    2 x num_key_value_heads 8 x head_dim 128 = 2,048 elements per token per layer everywhere; only the window differs. Past 4,096 tokens the marginal cost is 81.92 MB per 1,000 tokens, with 46.7 MB of fixed window state.

  • K-EXAONE 2.0 750B-A37B Attention mechanism official claim high confidence

    Model card: "1 x Global (NoPE) / 1 x 4096 SWA / 19 x [3 x 128 SWA + 1 x Global] Blocks", and the sliding_windows array matches exactly — 20 layers at 0 (full attention), one at 4,096 and 57 at 128.

  • K-EXAONE 2.0 750B-A37B Weights availability official claim high confidence
  • Nemotron 3 Ultra 550B-A55B Parameters, licence and weights official claim high confidence
  • Nemotron 3 Ultra 550B-A55B KV-cache geometry estimated high confidence

    Attention layers cache 2 x num_key_value_heads 2 x head_dim 128 = 512 elements per token, so the model adds only 12.29 MB per 1,000 tokens. The state-space layers hold mamba_num_heads 256 x mamba_head_dim 64 x ssm_state_size 128 = 2,097,152 elements at fp32 plus a convolution window of (256 x 64 + 2 x n_groups 8 x 128) x (conv_kernel 4 - 1) = 55,296 elements at bf16, giving 8,499,200 bytes per layer — about 408 MB per sequence before a single token of context. That is the layout both runtimes allocate for Mamba-2: vLLM e52be1a6's mamba2_state_shape (model_executor/layers/mamba/mamba_utils.py:195) and SGLang fa663e72's Mamba2StateShape.create (srt/configs/mamba_utils.py:195). CORRECTED 2026-09-13: the convolution state was 73,728 elements at fp32, 2,170,880 elements per layer and about 417 MB per sequence. A Mamba-2 window keeps conv_kernel - 1 = 3 inputs, and both runtimes store it at the model's bfloat16 while the state stays at float32: vLLM takes the window at the model dtype and the state from mamba_ssm_cache_dtype (_mamba_state_dtype, model_executor/layers/mamba/mamba_utils.py:98), which model_executor/models/config.py sets from this checkpoint's own mamba_ssm_cache_dtype; SGLang takes the window at SGLANG_MAMBA_CONV_DTYPE, default bfloat16 (srt/environ.py:1348), and the state at its float32 default. The split by head is exact while the tensor-parallel size divides both the 8 groups and the 256 heads: 1, 4 and 8 do and 72 does not, and vLLM's MambaMixer2 asserts both (model_executor/layers/mamba/mamba_mixer2.py:305 and :309, the second waived only for a single group). vLLM's --mamba-ssm-cache-dtype bfloat16 and SGLang's --mamba-ssm-dtype bfloat16 each halve the state.

  • Nemotron 3 Ultra 550B-A55B Attention mechanism official claim high confidence

    config.json layers_block_type over 108 entries: 48 mamba, 48 moe and 12 attention. NVIDIA warns of the consequence in its own report — below about 64K tokens the 32-bit Mamba state is larger than the FP8 KV cache, so the fixed term dominates rather than the growing one.

  • Nemotron 3 Ultra 550B-A55B Weights availability official claim high confidence
  • MiniMax-M3 Parameters, licence and weights official claim high confidence
  • MiniMax-M3 KV-cache geometry estimated high confidence

    config.json: 2 (key and value) x num_key_value_heads 4 x head_dim 128 = 1,024 elements per token per layer, across 60 layers — 120 KiB per token at bf16, which is more per token than either MLA model here despite being a much smaller model. That is the trade grouped-query attention makes. Sparse attention selects the top 16 blocks of 128 to *read*, which reduces arithmetic and not storage, so it does not enter this figure. Only 4 KV heads means the cache stops dividing beyond 4 accelerators and replicates thereafter. EXCLUDED, and worth knowing: MiniMax Sparse Attention runs an index branch that keeps its own key-only cache, allocated separately from the main K/V by both SGLang and vLLM. sparse_attention_config gives sparse_index_dim 128 and sparse_num_index_heads 4, and its sparse_attention_freq array marks 57 of the 60 layers as sparse — so the term is between 11.9% of the figure above (one shared 128-element index key per token per sparse layer) and 47.5% (one per index head). The release does not say which, and the gap between those readings is wider than several models in this table, so the figure above covers the main branch only. This is the same omission the GLM-5.2 entry records for its DeepSeek-style indexer; it was not recorded here, which was an inconsistency in treatment on the second most expensive model in the catalog.

  • MiniMax-M3 Attention mechanism official claim high confidence

    config.json: 64 query heads over 4 KV heads at head dimension 128, across 60 layers. sparse_attention_config selects the top 16 blocks of 128 on layers 3 to 59; layers 0 to 2 attend densely. Sparse *selection* does not shrink the cache — the entries have to exist to be selectable — so this is sized as an ordinary grouped-query cache, which is the honest reading and the conservative one.

  • MiniMax-M3 Expert routing official claim high confidence

    config.json: num_local_experts 128, num_experts_per_tok 4, n_shared_experts 1.

  • MiniMax-M3 Weights availability official claim high confidence
  • MiniMax-M3 Precision: MXFP8 (MiniMax) official claim high confidence

    safetensors index metadata.total_size 451,543,283,200 bytes.

  • MiniMax-M3 Precision: NVFP4 (NVIDIA) measured high confidence

    Size summed from the repository file listing: 250,103,762,320 bytes across 88 safetensors shards, header bytes included, because the index declares no metadata.total_size. The accuracy figures are NVIDIA's own evaluation of its build, and unlike the AtomicChat entries the repository publishes no evaluation logs, calibration corpus or reference logits alongside them. CORRECTED 2026-09-13: this note said they were reproducible from logs, a calibration corpus and reference logits published alongside them, the sentence the AtomicChat GGUF entries carry, where it is true.

  • Hunyuan Hy3 Parameters, licence and weights official claim high confidence
  • Hunyuan Hy3 KV-cache geometry estimated high confidence

    2 (key and value) x num_key_value_heads 8 x head_dim 128 = 2,048 elements per token per layer, across all 80 layers: 327.68 MB per 1,000 tokens at bf16. That is 26.7 times what Nemotron 3 Ultra costs at nearly twice the parameter count, and the clearest illustration in this catalog that long-context cost is set by attention design rather than by parameter count. CORRECTED 2026-08-21: this read 'forty times a Mamba hybrid of twice the size', which welded two different comparisons together — 40x is the ratio against Nemotron 3 Super, a model 0.4 times this size, not 2 times it.

  • Hunyuan Hy3 Attention mechanism official claim high confidence

    Model card: "Attention Heads | 64 (GQA, 8 KV heads, head dim 128)". No sliding window, no latent compression, no linear layers.

  • Hunyuan Hy3 Weights availability official claim high confidence
  • Command A+ (05-2026) Parameters, licence and weights official claim medium confidence

    Parameter counts and licence from the model card; the exact release day was not published, so the first of the month is recorded and the confidence lowered accordingly.

  • Command A+ (05-2026) KV-cache geometry estimated high confidence

    2 x num_key_value_heads 8 x head_dim 128 = 2,048 elements per token per layer throughout; the sliding layers stop at 4,096 tokens and contribute a fixed 402.7 MB per sequence thereafter. Past the window the marginal cost is 32.77 MB per 1,000 tokens against 131.07 MB inside it — a fourfold flattening.

  • Command A+ (05-2026) Attention mechanism official claim high confidence

    Model card: "The attention layers interleave sliding-window attention layers with Rotational Positional Embeddings and global attention layers without positional embeddings in a 3:1 ratio". config.json layer_types confirms 24 sliding and 8 full over 32 layers, sliding_window 4096.

  • Command A+ (05-2026) Expert routing official claim high confidence

    text_config: num_experts 128, num_experts_per_tok 8, num_shared_experts 4, expert_selection_fn sigmoid, shared_expert_combination_strategy average, first_k_dense_replace 0 - so every layer routes, with no dense prefix.

  • Command A+ (05-2026) Weights availability official claim high confidence
  • Nemotron 3 Super 120B-A12B Parameters, licence and weights official claim high confidence
  • Nemotron 3 Super 120B-A12B KV-cache geometry estimated high confidence

    8 attention layers at 2 x 2 x 128 = 512 elements per token give 8.19 MB per 1,000 tokens. Each state-space layer holds mamba_num_heads 128 x mamba_head_dim 64 x ssm_state_size 128 = 1,048,576 elements at fp32 plus a convolution window of (128 x 64 + 2 x n_groups 8 x 128) x (conv_kernel 4 - 1) = 30,720 elements at bf16, about 170 MB per sequence in total. That is the layout both runtimes allocate for Mamba-2: vLLM e52be1a6's mamba2_state_shape (model_executor/layers/mamba/mamba_utils.py:195) and SGLang fa663e72's Mamba2StateShape.create (srt/configs/mamba_utils.py:195). CORRECTED 2026-09-13: the convolution state was 40,960 elements at fp32, about 174 MB per sequence in total. A Mamba-2 window keeps conv_kernel - 1 = 3 inputs, and both runtimes store it at the model's bfloat16 while the state stays at float32: vLLM takes the window at the model dtype and the state from mamba_ssm_cache_dtype (_mamba_state_dtype, model_executor/layers/mamba/mamba_utils.py:98), which model_executor/models/config.py sets from this checkpoint's own mamba_ssm_cache_dtype; SGLang takes the window at SGLANG_MAMBA_CONV_DTYPE, default bfloat16 (srt/environ.py:1348), and the state at its float32 default. The split by head is exact while the tensor-parallel size divides both the 8 groups and the 128 heads: 1, 4 and 8 do and 72 does not, and vLLM's MambaMixer2 asserts both (model_executor/layers/mamba/mamba_mixer2.py:305 and :309, the second waived only for a single group). vLLM's --mamba-ssm-cache-dtype bfloat16 and SGLang's --mamba-ssm-dtype bfloat16 each halve the state.

  • Nemotron 3 Super 120B-A12B Attention mechanism official claim high confidence

    config.json hybrid_override_pattern resolves to 40 mamba, 40 moe and 8 attention layers of 88.

  • Nemotron 3 Super 120B-A12B Weights availability official claim high confidence
  • Mistral Large 3 Parameters, licence and weights official claim high confidence
  • Mistral Large 3 KV-cache geometry estimated high confidence

    params.json: kv_lora_rank 512 + qk_rope_head_dim 64 = 576 cached elements per token per layer across all 61 layers, or 68.6 KiB per token at bf16. As with every MLA model here the latent is a single KV head, so ordinary tensor parallelism replicates it on every rank while the weights divide.

  • Mistral Large 3 Attention mechanism official claim high confidence

    params.json: kv_lora_rank 512, qk_rope_head_dim 64, 61 layers. The n_kv_heads field reads 128 and is not the cache width — applying the grouped-query formula to it would overstate the cache by more than a hundredfold.

  • Mistral Large 3 Expert routing official claim high confidence

    params.json: num_experts 128, num_experts_per_tok 4, num_shared_experts 1.

  • Mistral Large 3 Weights availability official claim high confidence
  • Mistral Large 3 Precision: BF16 (training precision) official claim high confidence

    consolidated.safetensors.index.json metadata.total_size 1,351,982,353,920 bytes.

  • Mistral Large 3 Precision: NVFP4 (Mistral / Red Hat) official claim high confidence

    consolidated.safetensors.index.json metadata.total_size 403,125,497,052 bytes.

  • Gemma 4 E4B Parameters, licence and weights official claim high confidence
  • Gemma 4 E4B KV-cache geometry estimated high confidence

    Only 24 of the 42 layers store a cache at all: num_kv_shared_layers is 18, and those layers reuse an earlier layer's keys and values instead of keeping their own. Which 18 comes from the reference implementation config.json names: Gemma4ForConditionalGeneration sets first_kv_shared_layer_idx to num_hidden_layers minus num_kv_shared_layers, 42 - 18 = 24, and every layer from 24 upward reads stored keys and values rather than projecting its own. So the storing layers are 0 through 23, and config.json's layer_types array says what each of them is — five sliding_attention then one full_attention, repeating — which puts full attention at 5, 11, 17 and 23. That is 4 global at 2 x 2 x global_head_dim 512 = 2,048 elements per token and 20 sliding at 2 x 2 x head_dim 256 = 1,024, capped at sliding_window 512: 16.38 MB per 1,000 tokens with a 21 MB fixed term. EVERY LAYER STILL READS ONE: from layer 24 each layer reads the last storing layer of its own type, 23 for full attention and 22 for sliding (transformers 93ebf6b1 Gemma4Attention store_full_length_kv and shared_kv_states; vLLM e52be1a6 kv_sharing_target_layer_name; SGLang fa663e72 kv_shared_layer_index), so 7 layers read the 4 global caches and 35 read the 20 sliding windows. The groups record those as readingLayers, which multiply what a step reads and not what memory holds. CORRECTED 2026-09-13: both bandwidth searches charged each stored cache once. UPGRADED 2026-09-11 from an assumption. The previous note read "ASSUMPTION: which 18 layers share is not documented for Gemma 4, so the split between global and sliding among the storing layers follows the Gemma 3n convention that the top layers are the sharers". The convention was the right guess and no figure changes, but it was never a guess: both halves are stated, one in the config and one in the class the config names, and neither had been read. The attention record beside this one had already recorded the 5:1 layer pattern, so the entry was calling a split an assumption in one field and quoting its evidence in another.

  • Gemma 4 E4B Attention mechanism official claim high confidence

    config.json: 5 sliding : 1 full over 42 layers, sliding_window 512, num_kv_shared_layers 18. The shared-layer mechanism means those layers materialise no cache of their own. The cache itself is sized in the kvCache block, not here. CORRECTED 2026-09-11: this note gave the reason for that as "the global layers publish no KV head count (num_global_key_value_heads is null)". A null there is not an absent figure. The reference implementation gives every full_attention layer a head_dim override of global_head_dim 512 and adds a num_key_value_heads override only when the global field is set, so null leaves those layers on the top-level count of 2 — which is what the kvCache block had already sized them at, 2 x 2 x 512. Confidence rises from medium to high with that resolved. CORRECTED 2026-08-29: the second source was gemma-4-qat-hf, which resolves to google/gemma-4-31B-it-qat-w4a16-ct — a quantized build of a different model, and one that says nothing about layer types either way. The config.json quoted here is E4B's own, so the citation is now the repository it was read from.

  • Gemma 4 E4B Weights availability official claim high confidence
  • Gemma 4 E4B Precision: INT4 (Google QAT) official claim high confidence

    Single-file checkpoint of 11,513,861,828 bytes, re-verified against the HuggingFace API on 2026-08-16 after an internal consistency check flagged the ratio as an outlier. The bf16 reference was confirmed at the same time: 15,992,595,884 bytes over 7,996,156,490 bf16 parameters.

  • Gemma 4 31B Parameters, licence and weights official claim high confidence
  • Gemma 4 31B KV-cache geometry estimated high confidence

    config.json layer_types repeats 5 sliding to 1 full over 60 layers. Sliding layers cache 2 x num_key_value_heads 16 x head_dim 256 = 8,192 elements per token but stop at sliding_window 1,024, so past the window they contribute a fixed 838.9 MB per sequence rather than a slope. Global layers take a per-layer override: transformers builds per_layer_config with head_dim = global_head_dim 512 for every full_attention layer, and attention_k_eq_v additionally swaps num_key_value_heads for num_global_key_value_heads 4. They therefore cache 2 x 4 x 512 = 4,096 elements per token, or 81.92 MB per 1,000 tokens past the window against 901.12 MB inside it. CORRECTED 2026-08-21: this was recorded as 2,048 on the reading that attention_k_eq_v makes keys and values one stored tensor. It does not. In modeling_gemma4.py the flag only drops v_proj, after which value_states starts as the key projection's output and then diverges — k_norm plus RoPE on one, a scale-free v_norm on the other — and both are passed to past_key_values.update(). It saves a weight matrix and a matmul, not cache. Two vLLM issue threads reporting Gemma 4 page-size math independently quote 512 x 2 for the global layers. The figure was understated by exactly a factor of two. CROSS-CHECKED 2026-08-27 against Sebastian Raschka's independently published BF16 cache-per-token figures, which agree exactly with this catalog on Muse Glimmer 30B and Qwen3.6 27B and disagree here: counting every layer as growing, before the window cap, this geometry gives 880.00 KiB per token against his 840 KiB. The whole 40 KiB is the ten global layers. 840.0 is what the ten global layers give at 2,048 elements per token instead of 4,096, which is half this entry's current figure. CORRECTED 2026-08-28: this sentence used to attribute that halving to substituting the top-level head_dim 256 for global_head_dim 512, and to add that the same substitution is what this entry carried until the 2026-08-21 correction above. The second half is false and the first is unknowable. Two independent misreadings land on 2,048 — 4 heads × 256 with the K-and-V factor of two, or 4 × 512 without it — and the entry made the SECOND one: before a606d95 the note read “Global layers set attention_k_eq_v, meaning keys and values are one unified tensor, so they cache num_global_key_value_heads 4 × global_head_dim 512 = 2,048”, which reads the per-layer override correctly and drops the factor of two. Which of the two produced the published 840 is not decidable from the number, so this entry no longer claims to know. config.json for google/gemma-4-31b-it declares head_dim 256 and global_head_dim 512 side by side with num_global_key_value_heads 4, and transformers applies the 512 to every full_attention layer. Two independent readers making the same substitution is the argument for stating the override explicitly rather than assuming a model has one head dimension.

  • Gemma 4 31B Attention mechanism official claim high confidence

    config.json layer_types: a 5 sliding : 1 full pattern repeating over 60 layers, sliding_window 1024. Local and global layers have different geometry — 16 KV heads at head dimension 256 locally against 4 at 512 globally — so they must be sized separately.

  • Gemma 4 31B Weights availability official claim high confidence
  • Gemma 4 31B Precision: INT4 (Google QAT) official claim high confidence

    Single-file checkpoint of 23,265,352,448 bytes.

  • Qwen3.6-27B Parameters, licence and weights official claim high confidence
  • Qwen3.6-27B KV-cache geometry estimated high confidence

    16 growing layers at 2 x num_key_value_heads 4 x head_dim 256 = 2,048 elements per token: 65.54 MB per 1,000 tokens. Linear layers hold linear_num_value_heads 48 x 128 x 128 = 786,432 elements at fp32 plus a convolution window of (2 x linear_num_key_heads 16 x 128 + 48 x 128) x (linear_conv_kernel_dim 4 - 1) = 30,720 elements at bf16, 153.94 MB per sequence. The recurrent-state layout is the one both runtimes allocate for Gated DeltaNet — vLLM e52be1a6's gated_delta_net_state_shape (model_executor/layers/mamba/mamba_utils.py:274) and SGLang fa663e72's Mamba2StateShape.create with the key heads as its groups (srt/configs/qwen3_next.py:300) — rather than a figure published for this model; it shifts the constant term, not the slope. CROSS-CHECKED 2026-08-27 against Sebastian Raschka's independently published BF16 cache-per-token figures: this geometry gives 64.00 KiB per token against his 64 KiB. CORRECTED 2026-08-27: that sentence used to describe the basis as “counting every layer as growing, before the window cap”, which is boilerplate from the Muse Glimmer and Gemma 4 notes and is wrong here twice over. No group on this model carries a window, and 48 of its 64 layers are Gated DeltaNet with no per-token term at all. The 64.00 is the 16 full-attention layers alone, at 2 × 4 KV heads × 256 = 2,048 elements per token per layer, so on this entry the modelled figure and the uncapped one are the same number. An exact match. CORRECTED 2026-08-29: every claim on this model cited qwen36-github, github.com/QwenLM/Qwen3.6, which GitHub now answers with a 301 to QwenLM/Qwen3.8 because the repository was renamed for the next generation. A renamed repository keeps no copy of what it used to hold, so the citation did not break loudly — it silently served a different generation's documentation to anyone checking this entry's arithmetic. The source record is removed and the eight provenance blocks it backed now cite the model's own Hugging Face repository, whose config.json was re-read today and agrees with every figure here: 64 layers split 16 full_attention and 48 linear_attention, num_key_value_heads 4 at head_dim 256, linear_num_value_heads 48 at key and value head dim 128, linear_conv_kernel_dim 4. CORRECTED 2026-09-13: the convolution state was 40,960 elements at fp32 and the constant 158.86 MB per sequence. The window keeps linear_conv_kernel_dim - 1 = 3 inputs, not 4, and both runtimes store it at the model's bfloat16 while the state follows mamba_ssm_dtype, float32: vLLM takes the window at the model dtype and the state from mamba_ssm_cache_dtype (_mamba_state_dtype, model_executor/layers/mamba/mamba_utils.py:98), which model_executor/models/config.py sets from this checkpoint's mamba_ssm_dtype; SGLang takes the window at SGLANG_MAMBA_CONV_DTYPE, default bfloat16 (srt/environ.py:1348), and the state from the same field. The split by head is exact while the tensor-parallel size divides both the 16 key heads and the 48 value heads: 1, 4 and 8 do and 72 does not. vLLM's --mamba-ssm-cache-dtype bfloat16 and SGLang's --mamba-ssm-dtype bfloat16 each halve the state.

  • Qwen3.6-27B Attention mechanism official claim high confidence

    Architecturally identical to Qwen3.8-27B — same model type and every attention field matches: 64 layers with 48 linear and 16 full, 4 KV heads at head dimension 256.

  • Qwen3.6-27B Weights availability official claim high confidence
  • Qwen3.6-35B-A3B Parameters, licence and weights official claim high confidence
  • Qwen3.6-35B-A3B KV-cache geometry estimated high confidence

    10 growing layers at 2 x 2 x 256 = 1,024 elements per token: 20.48 MB per 1,000 tokens. Linear layers hold linear_num_value_heads 32 x 128 x 128 = 524,288 elements at fp32 plus a convolution window of (2 x linear_num_key_heads 16 x 128 + 32 x 128) x (linear_conv_kernel_dim 4 - 1) = 24,576 elements at bf16, about 64 MB per sequence. The convolution term is the layout both runtimes allocate for Gated DeltaNet — vLLM e52be1a6's gated_delta_net_state_shape (model_executor/layers/mamba/mamba_utils.py:274) and SGLang fa663e72's Mamba2StateShape.create with the key heads as its groups (srt/configs/qwen3_next.py:300) — rather than a figure published for this model. CORRECTED 2026-08-29: every claim on this model cited qwen36-github, github.com/QwenLM/Qwen3.6, which GitHub now answers with a 301 to QwenLM/Qwen3.8 because the repository was renamed for the next generation. A renamed repository keeps no copy of what it used to hold, so the citation did not break loudly — it silently served a different generation's documentation to anyone checking this entry's arithmetic. The source record is removed and the eight provenance blocks it backed now cite the model's own Hugging Face repository, whose config.json was re-read today and agrees with every figure here: 40 layers split 10 full_attention and 30 linear_attention, num_key_value_heads 2 at head_dim 256, linear_num_value_heads 32 at key and value head dim 128. CORRECTED 2026-09-13: the convolution state was 32,768 elements at fp32, about 67 MB per sequence. The window keeps linear_conv_kernel_dim - 1 = 3 inputs, not 4, and both runtimes store it at the model's bfloat16 while the state follows mamba_ssm_dtype, float32: vLLM takes the window at the model dtype and the state from mamba_ssm_cache_dtype (_mamba_state_dtype, model_executor/layers/mamba/mamba_utils.py:98), which model_executor/models/config.py sets from this checkpoint's mamba_ssm_dtype; SGLang takes the window at SGLANG_MAMBA_CONV_DTYPE, default bfloat16 (srt/environ.py:1348), and the state from the same field. The split by head is exact while the tensor-parallel size divides both the 16 key heads and the 32 value heads: 1, 4 and 8 do and 72 does not. vLLM's --mamba-ssm-cache-dtype bfloat16 and SGLang's --mamba-ssm-dtype bfloat16 each halve the state.

  • Qwen3.6-35B-A3B Attention mechanism official claim high confidence

    config.json: 40 layers, 30 linear and 10 full, 2 KV heads at head dimension 256, linear_num_value_heads 32.

  • Qwen3.6-35B-A3B Weights availability official claim high confidence
  • Qwen3.5-397B-A17B Parameters, licence and weights official claim high confidence

    Parameter count and weight size re-read 2026-08-29 from the repository's own API dtype split and safetensors index rather than from the model's name.

    Qwen (Hugging Face) ↗ read 2026-08-29
  • Qwen3.5-397B-A17B KV-cache geometry estimated high confidence

    15 growing layers at 2 x 2 x 256 = 1,024 elements per token: 30.72 MB per 1,000 tokens. The recurrent state is linear_num_value_heads 64 x 128 x 128 = 1,048,576 elements per layer at fp32, beside a convolution window of (2 x linear_num_key_heads 16 x 128 + 64 x 128) x (linear_conv_kernel_dim 4 - 1) = 36,864 elements at bf16: 4,268,032 bytes per layer, about 192.06 MB per sequence. That is the layout both runtimes allocate for Gated DeltaNet: vLLM e52be1a6's gated_delta_net_state_shape (model_executor/layers/mamba/mamba_utils.py:274) and SGLang fa663e72's Mamba2StateShape.create with the key heads as its groups (srt/configs/qwen3_next.py:300). CORRECTED 2026-08-21: this note previously excluded the short-convolution state on the grounds that "the key-head count it depends on is not published for this model". It is: config.json carries linear_num_key_heads 16 alongside linear_num_value_heads 64, linear_key_head_dim 128 and linear_conv_kernel_dim 4. Applying the same formula that reproduces every sibling entry exactly — (2 x 16 x 128 + 64 x 128) x 4 = 49,152 — the constant term rises from 188.74 MB to 197.59 MB, 4.7% higher. The slope is unaffected. CORRECTED 2026-09-13: the formula the correction of 2026-08-21 applied, (2 x 16 x 128 + 64 x 128) x 4, counted the window as the whole kernel of four inputs and priced it at fp32, and so did every sibling entry it reproduced; the constant was 1,097,728 elements per layer at fp32, 197.59 MB per sequence. The window keeps linear_conv_kernel_dim - 1 = 3 inputs, not 4, and both runtimes store it at the model's bfloat16 while the state follows mamba_ssm_dtype, float32: vLLM takes the window at the model dtype and the state from mamba_ssm_cache_dtype (_mamba_state_dtype, model_executor/layers/mamba/mamba_utils.py:98), which model_executor/models/config.py sets from this checkpoint's mamba_ssm_dtype; SGLang takes the window at SGLANG_MAMBA_CONV_DTYPE, default bfloat16 (srt/environ.py:1348), and the state from the same field. The split by head is exact while the tensor-parallel size divides both the 16 key heads and the 64 value heads: 1, 4 and 8 do and 72 does not. vLLM's --mamba-ssm-cache-dtype bfloat16 and SGLang's --mamba-ssm-dtype bfloat16 each halve the state.

  • Qwen3.5-397B-A17B Attention mechanism official claim medium confidence

    Model card: "15 * (3 * (Gated DeltaNet -> MoE) -> 1 * (Gated Attention -> MoE))" over 60 layers, 2 KV heads at head dimension 256, linear_num_value_heads 64.

  • Qwen3.5-397B-A17B Expert routing official claim high confidence

    text_config: num_experts 512, num_experts_per_tok 10, moe_intermediate_size 1024, shared_expert_intermediate_size 1024 over 60 layers at hidden_size 4096.

    Qwen (Hugging Face) ↗ read 2026-08-29
  • Qwen3.5-397B-A17B Weights availability official claim high confidence
  • DeepSeek V4 Flash Parameters, licence and weights official claim high confidence

    CONFLICT preserved: DeepSeek labels these checkpoints "FP4 + FP8 Mixed" while defaultWeightBits stores a single number. CORRECTED 2026-09-11: defaultWeightBits read 8 while nativeFormat.bits read 4, so the entry contradicted itself. The 8 read as a conservative capacity assumption and was not one — defaultWeightBits is consulted only when no checkpoint size is published, which is never the case here, so it reached no arithmetic at all and would have doubled the estimate the first time a DeepSeek entry arrived without a published size. It is now 4, which is also what the artefact measures: 0.588 bytes per parameter for Flash and 0.541 for Pro, against a range of 1.003 to 1.143 across the 12 eight-bit builds this catalog has measured. A schema rule now rejects the disagreement.

  • DeepSeek V4 Flash KV-cache geometry estimated high confidence

    Read from DeepSeek's own reference implementation (inference/model.py in the release repository), which settles three things config.json alone does not. (1) ONE ENTRY IS 512 ELEMENTS, NOT 576. Attention sets nope_head_dim = head_dim - rope_head_dim = 512 - 64 = 448 and rotates kv[..., -64:], so RoPE occupies the last 64 of the 512 rather than adding to them; the cache buffer is registered as [batch, size, head_dim] and a single wkv projection fills it, so there is no separate K and V either. (2) EVERY LAYER ALSO HOLDS A 128-TOKEN SLIDING WINDOW: kv_cache_size = window_size + max_seq_len // compress_ratio, so 43 x 128 x 512 elements sit under the whole model as a fixed 5.64 MB per sequence. (3) A LAYER AT RATIO 0 HAS NO COMPRESSED ENTRIES AT ALL — the first two layers are the window and nothing else, so they contribute no slope. compress_ratios in the 0731 checkpoint gives 2 layers at 0, 21 at 4 and 20 at 128 over the 43 modelled layers, so the slope is 21 x 128 + 20 x 4 = 2,768 elements per token: 5.54 MB per 1,000 tokens. CORRECTED 2026-08-21: this entry read 576 elements per entry with no window term and an unbounded slope on the two ratio-0 layers, giving 8.53 MB per 1,000 tokens — 1.54x the real figure. The 576 was DeepSeek V3's concatenated kv_lora_rank + qk_rope_head_dim carried across to an architecture that dropped both fields. STILL EXCLUDED: the sparse indexer's own key cache. index_head_dim is 128 and only the 21 ratio-4 layers run an indexer. The reference code keeps one 128-element key per compressed position per layer, shared by all 64 index heads: Indexer registers kv_cache as [max_batch_size, max_seq_len // compress_ratio, index_head_dim] and scores every head's query against that one key (inference/model.py:395, 405, 426). CORRECTED 2026-09-13: this said the release does not state whether the keys are stored per head (index_n_heads 64) or shared. The window is modelled as a ramp capped at 128 tokens rather than as a flat constant, so a sequence shorter than the window is not charged for the whole of it — the same way Gemma 4's window is expressed, which is what makes the two entries comparable on their constant terms.

  • DeepSeek V4 Flash Attention mechanism official claim high confidence

    Technical report section 4.2.1: 43 layers, first two pure sliding-window, the rest CSA and HCA "in an interleaved manner". CSA compression rate 4 with top-k 512; HCA compression rate 128; sliding window 128; 64 query heads at head dimension 512. The report never states how many layers are of each kind, and gives no per-token cache figure, so the cache is not sized here. It does say the uncompressed sliding-window entries "exist in every layer" and run "approximately 8 times larger than the compressed CSA and HCA KV entries" — which means the window, not the compressed part, dominates. WHAT A STEP READS is in the release's own inference/model.py, not the report. index_topk counts compressed entries: the Indexer keeps topk(min(index_topk, end_pos // ratio)) of the compressed index keys (line 433), so a ratio-4 layer reads at most 512 entries of four tokens each, 2,048 tokens' worth, which is that group's tokensPerEntry. A ratio-128 layer builds no indexer (lines 474-477), and get_compress_topk_idxs hands it every compressed position (lines 275-282, called at 519), so that group reads every entry. vLLM e52be1a6 and SGLang fa663e72 select the same way: index_topk entries on ratio-4 layers and every compressed entry on ratio-128 layers. CORRECTED 2026-09-13: selectedEntries was selectedTokens, and the bandwidth lower bound charged the ratio-4 layers 512 tokens rather than 512 entries and capped the ratio-128 layers at 512 tokens as well.

  • DeepSeek V4 Flash Expert routing official claim high confidence

    config.json: n_routed_experts 256, num_experts_per_tok 6, n_shared_experts 1.

  • DeepSeek V4 Flash Weights availability official claim high confidence
  • DeepSeek V4 Flash Precision: IQ3_M (AtomicChat) measured high confidence

    Size verified to the byte against the repository listing. Accuracy is the quantizer’s own published KL-divergence and top-1 measurement, with per-file logs.

  • DeepSeek V4 Pro Parameters, licence and weights official claim medium confidence

    CONFLICT preserved on the active parameter count: 49B is the preview figure and DeepSeek has not republished it for the 0813 build, while one third-party redistribution says 48B. Total parameters and checkpoint size are solid; the active count is inherited. CORRECTED 2026-09-11: defaultWeightBits read 8 while nativeFormat.bits read 4, so the entry contradicted itself. The 8 read as a conservative capacity assumption and was not one — defaultWeightBits is consulted only when no checkpoint size is published, which is never the case here, so it reached no arithmetic at all and would have doubled the estimate the first time a DeepSeek entry arrived without a published size. It is now 4, which is also what the artefact measures: 0.588 bytes per parameter for Flash and 0.541 for Pro, against a range of 1.003 to 1.143 across the 12 eight-bit builds this catalog has measured. A schema rule now rejects the disagreement.

  • DeepSeek V4 Pro KV-cache geometry estimated high confidence

    Same derivation as V4 Flash, from the same reference implementation: one entry is head_dim 512 elements with RoPE occupying its last 64, not 512 + 64, and every layer carries a 128-token sliding window on top of its compressed entries. compress_ratios in the 0813 checkpoint gives 30 layers at ratio 4 and 31 at 128 across the 61 modelled layers — no ratio-0 layers here, unlike Flash — so the slope is 30 x 128 + 31 x 4 = 3,964 elements per token, 7.93 MB per 1,000 tokens, over a fixed 8.00 MB per sequence of window. CORRECTED 2026-08-21: recorded as 8.92 MB per 1,000 tokens with no window term, on the same 576-element reading that Flash carried. Same exclusion: the indexer's own key cache, which the reference code keeps as one 128-element key per compressed position on each of the 30 ratio-4 layers, shared by all 64 index heads (inference/model.py:395, 405, 426, byte-identical to V4 Flash's). CORRECTED 2026-09-13: this said the release does not state its storage layout. The window is modelled as a ramp capped at 128 tokens rather than as a flat constant, so a sequence shorter than the window is not charged for the whole of it — the same way Gemma 4's window is expressed, which is what makes the two entries comparable on their constant terms.

  • DeepSeek V4 Pro Attention mechanism official claim high confidence

    Technical report section 4.2.1: 61 layers, first two HCA, the rest CSA and HCA interleaved. CSA compression rate 4 with top-k 1,024; HCA compression rate 128; sliding window 128; 128 query heads at head dimension 512. The layer split is not published, so the cache is not sized here. WHAT A STEP READS is in the release's inference/model.py, which is byte-identical to V4 Flash's, where the lines are cited: index_topk counts compressed entries, so a ratio-4 layer reads at most 1,024 entries of four tokens each, 4,096 tokens' worth, and a ratio-128 layer builds no indexer and reads every compressed entry. CORRECTED 2026-09-13: selectedEntries was selectedTokens, and the bandwidth lower bound charged every layer 1,024 tokens.

  • DeepSeek V4 Pro Expert routing official claim high confidence

    config.json: n_routed_experts 384, num_experts_per_tok 6, n_shared_experts 1.

  • DeepSeek V4 Pro Weights availability official claim high confidence
  • MiMo V2.5 Parameters, licence and weights official claim high confidence
  • MiMo V2.5 KV-cache geometry estimated high confidence

    Key and value head widths differ here, so the usual doubling does not apply: it is kv_heads x (head_dim + v_head_dim). Global layers give 4 x (192 + 128) = 1,280 elements per token; sliding layers give 8 x 320 = 2,560, capped at 128 tokens for a fixed 25.6 MB per sequence. Long-context marginal cost is 23.04 MB per 1,000 tokens.

  • MiMo V2.5 Attention mechanism official claim medium confidence

    Model card: "interleaving Sliding Window Attention (SWA) and Global Attention (GA) with a 5:1 ratio and 128 sliding window. This reduces KV-cache storage by nearly 6x." The hybrid_layer_pattern in config.json actually resolves to 9 global and 39 sliding of 48, which is 4.3:1 rather than 5:1. CONFLICT: the card states 8 KV heads for global and 4 for sliding, and the config states the reverse; the config is used here, and the card itself carries a notice that config.json was corrected after release.

  • MiMo V2.5 Weights availability official claim high confidence
  • MiMo V2.5 Pro Parameters, licence and weights official claim high confidence
  • MiMo V2.5 Pro KV-cache geometry estimated high confidence

    kv_heads 8 x (head_dim 192 + v_head_dim 128) = 2,560 elements per token on both layer kinds; only the window differs. 51.20 MB per 1,000 tokens at long context, with 39.32 MB of fixed window state.

  • MiMo V2.5 Pro Attention mechanism official claim medium confidence

    config.json hybrid_layer_pattern resolves to 10 global and 60 sliding layers of 70, sliding_window 128, with 8 KV heads and split key/value head dimensions of 192 and 128 throughout.

  • MiMo V2.5 Pro Weights availability official claim high confidence
  • GLM-5.3 Parameters, licence and weights official claim high confidence

    Read from the artefact rather than inherited, since 2026-08-29. The active-parameter count is the one figure still not from Z.ai: they have never published one for either version, and 40B is Artificial Analysis's.

  • GLM-5.3 KV-cache geometry estimated high confidence

    config.json: 78 layers, kv_lora_rank 512 plus qk_rope_head_dim 64 for a 576-element MLA latent, num_nextn_predict_layers 1. Previously inherited from GLM-5.2 and now read from GLM-5.3's own config, which agrees with it. DOWNGRADED 2026-08-29: this was the only one of the catalog's 33 KV geometries graded official-claim, and it is now estimated like the other 32. Z.ai publishes no per-token cache figure for GLM-5.3. What the config publishes is kv_lora_rank and qk_rope_head_dim; adding them, multiplying by 78 layers and calling the product a cache size is this repository's arithmetic, and the GLM-5.2 entry, which performs that same arithmetic on those same two fields, has always said estimated. EXCLUDED: the DSA lightning indexer keeps its own key cache and is not priced here, the same omission GLM-5.2 and Hunyuan HY4 each declare. Of the 78 layers modelled above, the 21 that indexer_types marks full keep one, and the 57 marked shared keep none and reuse the top-k of the full layer before them. vLLM e52be1a6 reads the same 21 off index_topk_freq 4 and index_skip_topk_offset 3 and builds an indexer only on those or on an MTP layer (models/deepseek_v32/attention.py:166-200); SGLang fa663e72 builds none on a layer dsa_layer_skips_topk marks shared and gives that layer a zero-row key buffer (srt/configs/model_config.py:311-319, srt/models/deepseek_v2.py:1808-1841, srt/mem_cache/index_key_cache.py:40-43). At index_head_dim 128 that is 21 x 128 x 2 bytes = 5.25 KiB per token at bf16, 6.0% on top of the modelled 87.75 KiB. The MTP layer (num_nextn_predict_layers 1) is not one of the 78: both runtimes build it only as the draft model for MTP speculative decoding (vLLM config/speculative.py:666-681, SGLang srt/configs/model_config.py:789-790), with a latent cache and an indexer key cache of its own, 576 + 128 elements, 1,408 bytes per token at bf16, which this entry does not count any more than it counts the drafting. CORRECTED 2026-09-13: this said the number of distinct indexer key caches is between 22 and 78 and that the config does not say which, computed its floor from 21 all the same, and put the ceiling at nearly four times that. The 22 counted the MTP layer, which the 78-layer latent leaves out, and neither runtime keeps a key cache on a shared layer, so 78 is no ceiling. The declaration was here while this entry inherited GLM-5.2's geometry and went missing when the geometry was re-derived from GLM-5.3's own config, which is how one edit raised the stated confidence and dropped the caveat that qualified it.

    Z.ai (Hugging Face) ↗ read 2026-08-29
  • GLM-5.3 Attention mechanism estimated medium confidence

    Inherited from GLM-5.2 on the vendor’s same-base-model statement.

  • GLM-5.3 Expert routing official claim high confidence

    config.json: n_routed_experts 256, num_experts_per_tok 8, n_shared_experts 1, first_k_dense_replace 3. Identical to GLM-5.2, as the vendor said it would be.

    Z.ai (Hugging Face) ↗ read 2026-08-29
  • GLM-5.3 Weights availability official claim high confidence

    The BF16 repository was last modified 2026-08-28 and the FP8 one 2026-08-29T09:51:13Z, both ungated. Z.ai's announcement on 2026-08-14 promised the weights “in two weeks after launch, once safety evaluation and hardening are complete”, which put the target at 2026-08-28. It landed on the 28th and the 29th.

  • GLM-5.3 Precision: BF16 official claim high confidence

    metadata.total_size 1,506,659,919,872 bytes over 59,585 tensors, byte-identical to GLM-5.2's index.

    Z.ai (Hugging Face) ↗ read 2026-08-29
  • GLM-5.3 Precision: NVFP4 (projected) estimated low confidence

    Two inferences stacked: that GLM-5.3 shares GLM-5.2’s architecture, and that somebody will quantize it the way NVIDIA quantized GLM-5.2. Carried because a factor of three in memory decides which hardware is even a candidate, and a stated projection is more useful than a blank — but it is the weakest figure in this catalog and is graded to say so.

  • GLM-5.2 Parameters, licence and weights official claim high confidence

    CONFLICT preserved: two independent reads of the bf16 checkpoint gave 1,506,659,919,872 bytes (the safetensors index’s declared total) and 1,506,689,458,421 bytes (the sum of the shards on disk). The declared total is carried, because tensor bytes rather than file bytes are what has to fit in accelerator memory; the difference is 0.002%.

  • GLM-5.2 KV-cache geometry estimated high confidence

    Derived from config.json: kv_lora_rank 512 + qk_rope_head_dim 64 = 576 cached elements per token per layer, across all 78 layers, which is 87.8 KiB per token at bf16. Two independent passes agreed on the fields. NOTE the trap: num_key_value_heads reads 64, but that is a compatibility field equal to num_attention_heads — applying the ordinary grouped-query formula to it would overstate this cache by roughly fifty times. EXCLUDED: the DSA lightning indexer keeps its own key cache, roughly 5 KiB per token (index_head_dim 128 over about twenty indexer layers at index_topk_freq 4), because how it shards under tensor parallelism is not published and a wrong sharding claim would be a larger error than the omission. The figures here are therefore about 5% low.

  • GLM-5.2 Attention mechanism official claim high confidence

    Z.ai states the consequence plainly: "Although the new GLM-5.2 architecture reduces per-token computational FLOPs, it does not proportionally reduce per-token KV-cache size."

  • GLM-5.2 Expert routing official claim high confidence

    config.json: n_routed_experts 256, num_experts_per_tok 8, n_shared_experts 1, first_k_dense_replace 3 (layers 0-2 are dense).

  • GLM-5.2 Weights availability official claim high confidence
  • GLM-5.2 Precision: FP8 (Z.ai) official claim high confidence

    safetensors index metadata.total_size 755,617,140,416 bytes.

  • GLM-5.2 Precision: NVFP4 (NVIDIA) measured high confidence

    Size from the safetensors index's metadata.total_size: 464,795,267,072 bytes. The accuracy figures are NVIDIA's own evaluation of its build against Z.ai's GLM-5.2-FP8, and unlike the AtomicChat entries the repository publishes no evaluation logs, calibration corpus or reference logits alongside them. CORRECTED 2026-09-13: the size was 464.82 GB, the 464,823,042,096 bytes of the repository's 47 files, which include each file's header as well as its tensors; and this note said the accuracy figures were reproducible from logs, a calibration corpus and reference logits published alongside them, the sentence the AtomicChat GGUF entries carry, where it is true.

  • Kimi K2.6 Parameters, licence and weights official claim high confidence
  • Kimi K2.6 KV-cache geometry estimated high confidence

    Derived from text_config: num_hidden_layers 61, and every layer caches kv_lora_rank 512 + qk_rope_head_dim 64 = 576 elements per token. num_nextn_predict_layers is 0, so unlike GLM-5.2 the safetensors index carries exactly 61 layer indices and there is no extra MTP layer to account for. num_key_value_heads reads 64, which on a latent-attention model does not describe the cache at all - it is the count of query-side heads the absorbed latent serves, and reading a KV size off it would overstate this model's cache by roughly 22x. The latent is a single KV head, so plain tensor parallelism replicates it on every rank; only decode context parallelism shards it. defaultBytesPerElement is 2 because quantization_config.kv_cache_scheme is null: the release prescribes no fp8 cache, so bf16 is the reference store, which is the rule this catalog applies everywhere else. fp8 remains available as a runtime lever.

  • Kimi K2.6 Attention mechanism official claim high confidence

    text_config: kv_lora_rank 512, q_lora_rank 1536, qk_nope_head_dim 128, qk_rope_head_dim 64, v_head_dim 128, num_attention_heads 64.

  • Kimi K2.6 Expert routing official claim high confidence

    text_config: n_routed_experts 384, num_experts_per_tok 8, n_shared_experts 1.

  • Kimi K2.6 Weights availability official claim high confidence
  • GLM-5.3-Flash Parameters, licence and weights official claim high confidence

    Total parameters are the repository's own safetensors count, 321,323,031,390, rather than the rounded 320B the model card prints — the same convention the glm-5-2 entry uses. Active parameters 18B is Z.ai's own published figure, not an aggregator's. Weight size is metadata.total_size for the default FP8 repository, 328,326,771,576 bytes, corroborated independently by vLLM's recipe at about 306 GiB. Context read from config.json rather than from the docs page's "1M".

  • GLM-5.3-Flash KV-cache geometry estimated high confidence

    Derived from config.json and modeling_glm5_next.py. layer_types splits 45 layers into 34 linear_attention and 11 deepseek_sparse_attention, and kda_layers lists exactly the 34. The latent is kv_lora_rank 512 plus qk_rope_head_dim 0 = 512 elements per token, not GLM-5.2's 576: mla_use_nope is true and the k_rot slice is empty, so copying the sibling entry's figure would overstate this model by 12.5%. The indexer stores 128 key elements, 128 gate scores and one validity flag = 257 elements per token per sparse-attention layer, from a single shared key head, which is why that group is replicated rather than by-kv-head: there is nothing to split across tensor-parallel ranks, so at TP=8 it costs eight times what a per-rank division would suggest. index_kpool 4 pools at scoring time out of the full stored tensor and does not divide the storage. Each linear layer holds linear_attn_config num_heads 64 x head_dim 128 x 128 = 1,048,576 state elements, which the code casts to fp32 before storing them, plus three convolution windows of 64 x 128 x (short_conv_kernel_size 4 - 1) = 24,576 elements each at bf16. Per-token slope 11 x 512 + 11 x 257 = 8,459 elements, 16.92 MB per 1,000 tokens at bf16; per-sequence constant 34 x (1,048,576 x 4 + 3 x 24,576 x 2) bytes = 147.62 MB per sequence. That per-token slope is 5.31x smaller than the glm-5-2 entry's, which is the quantitative content of Z.ai's claim to cut long-context serving cost. NOT VALIDATED against a published block size: unlike Kimi K3 no per-token or per-block figure is published for this model, and vLLM's KV-pool token counts cannot be converted to bytes per token without assumptions that move the answer threefold. The fp8 cache path is real but hardware-gated: vLLM's own recipe states Hopper must run a bf16 KV cache for this model, so the fp8 control here is reachable on Blackwell and newer only. CORRECTED 2026-09-13: the windows were 32,768 elements each at fp32, 1,146,880 elements per layer and 155.98 MB per sequence, and the fp32 cast was stated of all of it. It is true of the state and only of the state: modeling_glm5_next.py at transformers 415e6d2f casts last_recurrent_state to torch.float32 at line 729 before update_recurrent_state, while the convolution window goes through causal_conv1d_update and update_conv_state (lines 662-674) with no cast. Both runtimes keep that window for short_conv_kernel_size - 1 = 3 inputs at the activation width and the state at float32. vLLM e52be1a6's Glm5Next KDA layer calls kda_state_dtype with the model dtype and no SSM override (models/glm5next/nvidia/kda.py:136), so --mamba-ssm-cache-dtype does not reach this model's state there, and kda_state_shape (model_executor/layers/mamba/mamba_utils.py:298) sizes the window; SGLang fa663e72 builds KimiLinearStateShape from linear_attn_config (srt/configs/glm5_next.py:265) with the window at SGLANG_MAMBA_CONV_DTYPE, default bfloat16, and its --mamba-ssm-dtype bfloat16 halves the state. The split by head is exact while the tensor-parallel size divides the 64 heads: 1, 4 and 8 do and 72 does not.

  • GLM-5.3-Flash Attention mechanism official claim high confidence

    config.json layer_types: 34 linear_attention, 11 deepseek_sparse_attention, 45 total.

  • GLM-5.3-Flash Expert routing official claim high confidence

    config.json: n_routed_experts 288, num_experts_per_tok 8, n_shared_experts 1, first_k_dense_replace 3, moe_intermediate_size 2048.

  • GLM-5.3-Flash Weights availability official claim high confidence
  • GLM-5.3-Flash Precision: BF16 official claim high confidence

    metadata.total_size 642,646,653,816 bytes.

  • Qwen3.8-Flash-Next Parameters, licence and weights official claim high confidence

    Weight size is metadata.total_size for the default BF16 repository, 359,999,963,128 bytes — declared as a JSON float, alone in this catalog, which a strict integer parser rejects. It reconciles exactly with the dtype breakdown: 179,999,981,424 BF16 parameters x 2 plus 35 I64 x 8. Context read from config.json, which sets max_position_embeddings to 262,144; the served qwen3.8-flash API advertises 1M as a production feature added on top of this checkpoint, and the two are not the same artefact.

  • Qwen3.8-Flash-Next KV-cache geometry estimated high confidence

    Derived from config.json and modeling_qwen4_exp.py. layer_types splits 48 layers into 36 linear_attention and 12 full_attention, consistent with full_attention_interval 4. Each full-attention layer caches 2 x num_key_value_heads 2 x head_dim 256 = 1,024 elements per token. The sparse indexer keeps its own key-only cache of indexer_kv_heads 1 x indexer_head_dim 128 = 128 elements per token on those same layers; it is a single shared key head, so it replicates under tensor parallelism rather than dividing, and indexer_compress_ratio 4 divides the scoring rather than the storage. Each linear layer holds linear_num_value_heads 48 x linear_value_head_dim 128 x linear_key_head_dim 128 = 786,432 state elements at fp32 per mamba_ssm_dtype, plus a convolution window over q and k at the key-head count and v at the value-head count, (16 x 128 x 2 + 48 x 128) x (linear_conv_kernel_dim 4 - 1) = 30,720 elements at bf16. Note the convolution convention differs from GLM's, which covers all three projections at one head count. Per-token slope 12 x 1,024 + 12 x 128 = 13,824 elements, 27.65 MB per 1,000 tokens at bf16. The per-layer embedding (ple_layer_ids [2]) runs a short convolution with a window of its own: hidden_size 2560 x hc_count 4 = 10,240 channels over (ple_conv_kernel_size 4 - 1) x ngram_size 3 = 9 positions, 92,160 elements at bf16, which neither runtime divides across tensor-parallel ranks. vLLM e52be1a6 marks its MambaSpec tp_replicated (models/qwen4_exp/nvidia/model.py:828-833, ple_layer.py:175), and SGLang fa663e72 sizes it without the tensor-parallel size (srt/configs/qwen4_exp.py:96-101) in a ShortConvPool beside every state slot (srt/mem_cache/memory_pool.py:1318). Beside that SGLang keeps the last ngram_size - 1 = 2 token ids per slot, 16 bytes, which rounds to nothing here. Per-sequence constant 36 x (786,432 x 4 + 30,720 x 2) + 92,160 x 2 bytes = 115.64 MB per sequence on one accelerator. No published per-token figure exists to validate this against. CORRECTED 2026-09-13: the window was 40,960 elements at fp32, 827,392 elements per layer and 119.14 MB per sequence, and the per-layer embedding's window was not recorded at all. The window keeps linear_conv_kernel_dim - 1 = 3 inputs, not 4, and both runtimes store it at the model's bfloat16 while the state follows mamba_ssm_dtype, float32: vLLM applies Qwen3.5's hybrid-cache contract to this architecture (Qwen4ExpForConditionalGenerationConfig, model_executor/models/config.py:858), which sets mamba_ssm_cache_dtype from that field and leaves the window at the model dtype (_mamba_state_dtype, model_executor/layers/mamba/mamba_utils.py:98); SGLang sizes it through Qwen3Next's Mamba2StateShape.create (srt/configs/qwen3_next.py:300) at SGLANG_MAMBA_CONV_DTYPE, default bfloat16. The split by head is exact while the tensor-parallel size divides both the 16 key heads and the 48 value heads: 1, 4 and 8 do and 72 does not. vLLM's --mamba-ssm-cache-dtype bfloat16 and SGLang's --mamba-ssm-dtype bfloat16 each halve the state.

  • Qwen3.8-Flash-Next Attention mechanism official claim high confidence

    config.json layer_types: 36 linear_attention, 12 full_attention, 48 total, with full_attention_interval 4.

  • Qwen3.8-Flash-Next Expert routing official claim high confidence

    config.json: num_experts 512 — Qwen's key name, not n_routed_experts — num_experts_per_tok 10, moe_intermediate_size 640, shared_expert_intermediate_size 640.

  • Qwen3.8-Flash-Next Weights availability official claim high confidence
  • Qwen3.8-Flash-Next Precision: FP8 (Qwen) official claim high confidence

    metadata.total_size 185,502,232,570 bytes, read through /resolve/main/ because the /raw/ endpoint returns a git-LFS pointer for an index this large.

  • Hunyuan Hy4-preview Parameters, licence and weights official claim high confidence

    Parameter count from the safetensors API rather than from the card, so it is the sum of the dtype split (BF16 779,930,197,312 plus F32 30,795,421) rather than a rounded headline. Active parameters are the card's 49B, which excludes the MTP layer's 0.7B and could not be derived from config.json alone.

  • Hunyuan Hy4-preview KV-cache geometry estimated high confidence

    kv_lora_rank 512 plus qk_rope_head_dim 64 gives 576 elements per token per layer across all 78 layers: 89.86 MB per 1,000 tokens at bf16, against 327.68 MB for Hy3 at 295B. The same geometry as GLM-5.3, which also has 78 MLA layers at 576. What this does not price: the sparse-attention indexer. Its keys are cached too, but only where an indexer runs: vLLM e52be1a6 builds one on the 21 layers indexer_types marks "full" and on the MTP layer (num_nextn_predict_layers 1), and none on the 57 marked "shared", which reuse the top-k of the full layer before them (models/hy_v4/nvidia/attention.py:321-331). Each keeps 128 fp8 elements and one fp32 scale per token (145, 171, 177-182). That is enough to price the cache, which this entry still does not. CORRECTED 2026-09-13: this said the number of distinct indexer caches was somewhere between 22 and 78 and that the config alone does not say which. The same choice was made for GLM-5.3, which has the same DeepSeek-style indexer and the same 21-against-57 split: the latent is modelled, the indexer is declared and not guessed. GLM-5.3-Flash and Qwen3.8-Flash-Next are the two entries that do price an indexer, because their runtimes document the geometry. The roadmap carries this as an open item; the figure here is therefore a floor, not a total. CORRECTED 2026-08-29: that sentence read "the indexer is recorded" and pointed at an entry where it was not. GLM-5.2 declares the omission; GLM-5.3 lost the declaration when its geometry was re-derived from its own config, so this note vouched for a sibling that had gone silent. The declaration is restored on GLM-5.3 in the same change rather than the claim being trimmed to fit what was there. That sentence read "GLM-5.3-Flash is the one entry" and was wrong on the day it was written: both index groups landed on 2026-08-27 and this entry was added on 2026-08-29, so the count was two before the claim existed. Counting a population from memory rather than from the file is what produced both errors in this note.

  • Hunyuan Hy4-preview Attention mechanism official claim high confidence

    config.json: layer_types is 78 entries of "deepseek_sparse_attention", use_dsa true, kv_lora_rank 512 plus qk_rope_head_dim 64 for a 576-element MLA latent, index_head_dim 128, index_n_heads 32, index_topk 2048, and indexer_types splits 21 "full" against 57 "shared". index_topk counts tokens, and every layer reads that many from its own latent cache: vLLM e52be1a6 builds the Indexer with topk_tokens = index_topk (models/hy_v4/nvidia/attention.py:143) over a top-k buffer of index_topk token positions (model.py:226-233); a shared layer builds no indexer and attends with the indices of the closest preceding full layer (attention.py:54-59, 321-331); and transformers 415e6d2f writes each layer's own keys and values to the cache before choosing its top-k or taking the previous full layer's (models/hy_v4/modeling_hy_v4.py:396-397, 455-472). So selectedEntries is 2,048 on all 78 layers. CORRECTED 2026-09-13: the entry was marked selective with no size, so the bandwidth lower bound charged these layers no cache at all.

  • Hunyuan Hy4-preview Weights availability official claim high confidence
  • Hunyuan Hy4-preview Precision: MXFP8 (Tencent) estimated medium confidence

    Estimated rather than official-claim because the repository publishes no size: metadata.total_size is absent from the index and 813.25 GB is derived from Hugging Face's per-dtype element counts (BF16 9,640,893,760, F32 30,795,421, F8_E4M3 770,289,303,552, U8 23,555,211,264) at their known widths. Confidence is medium because a second method, summing the shards on disk, disagrees by 0.06%. Format is MXFP8 rather than plain E4M3, read off quant_algo and confirmed in the tensor shapes: down_proj [256,6144,2048] against down_proj_scale [256,6144,64] is a block of 32.

  • Granite 4.2 30B Parameters, licence and weights official claim high confidence

    29.28B is the safetensors API's single-dtype BF16 count of 29,276,770,304, not the card's rounded “30B”. Dense, so every parameter is active and no separate active count is carried.

    IBM (Hugging Face) ↗ read 2026-08-29
  • Granite 4.2 30B KV-cache geometry estimated high confidence

    2 (key and value) × num_key_value_heads 8 × head_dim 128 = 2,048 elements per token per layer across all 64 layers: 262.14 MB per 1,000 tokens at bf16. For the same size class that is 262.14 against 53.25 for Muse Glimmer 30B and 65.54 for Qwen3.8-27B, and 80% of what 295B Hunyuan Hy3 costs. Filling the 131,072-token window for a single sequence takes 34.4 GB of cache — more than the weights themselves at fp8.

  • Granite 4.2 30B Attention mechanism official claim high confidence

    config.json: num_hidden_layers 64, num_attention_heads 32, num_key_value_heads 8 for 4:1 grouping, and no head_dim key, so the head dimension is hidden_size / num_attention_heads = 4096 / 32 = 128. No sliding_window, no layer_types, no kv_lora_rank.

  • Granite 4.2 30B Weights availability official claim high confidence
  • Ling-3.0-flash Parameters, licence and weights official claim high confidence

    Parameter count from the safetensors API rather than the card, so it is the sum of the dtype split (BF16 127,486,240,128 plus F32 165,472) rather than a rounded headline. Active parameters are the card's 5.1B, which config.json alone does not give.

  • Ling-3.0-flash KV-cache geometry estimated high confidence

    kv_lora_rank 512 plus qk_rope_head_dim 64 gives 576 elements per token per layer on the seven latent layers: 8.06 MB per 1,000 tokens at bf16. Three entries here are cheaper still, at 5.54, 6.14 and 7.93, and every one of them buys it the same way: by keeping most of its layers off the per-token cache. The 35 linear layers contribute no slope at all. Their state is what fla.ops.kda returns, shape [N, HV, K, V] with HV = num_attention_heads 32 and K = V = head_dim 128, so 524,288 elements, and chunk.py asserts it must be float32. Beside it sit the three short convolutions for q, k and v, each over D = 32 × 128 = 4,096 channels. The ShortConvolution cache in the reference code is [N, D, W] with W = short_conv_kernel_size 4, but the runtimes keep one input fewer: SGLang fa663e72, which this layout names, builds KimiLinearStateShape from num_attention_heads, head_dim and short_conv_kernel_size (srt/configs/bailing_hybrid.py:208) and holds 3 × 4,096 × (4 - 1) = 36,864 elements at SGLANG_MAMBA_CONV_DTYPE, default bf16. The group records the two as separate components, each at its own width: 524,288 × 4 + 36,864 × 2 = 2,170,880 bytes per layer per sequence, or 75.98 MB per sequence across the 35 layers. That constant is why this model is not automatically the cheapest choice at short context, only at long. CORRECTED 2026-09-13: the windows were 49,152 elements, the full W = 4 of the reference cache, and the constant 76.84 MB per sequence. vLLM e52be1a6 keeps the same three inputs (kda_state_shape, model_executor/layers/mamba/mamba_utils.py:298, called from model_executor/models/bailing_moe_v3.py:618) at the model dtype, and its BailingMoeV3 layer calls kda_state_dtype without an SSM override (model_executor/models/bailing_moe_v3.py:611), so --mamba-ssm-cache-dtype does not reach this model's state there; SGLang's --mamba-ssm-dtype bfloat16 halves it. The split by head is exact while the tensor-parallel size divides the 32 heads: 1, 4 and 8 do and 72 does not.

  • Ling-3.0-flash Attention mechanism official claim high confidence

    The split is read off the model's own code, not inferred from the name. modeling_bailing_moe_v3.py lines 1004-1014 select BailingMoeV3MultiLatentAttention when (layer_idx + 1) % layer_group_size == 0 and BailingMoeV3KimiDeltaAttention otherwise; with num_hidden_layers 42 and layer_group_size 6 that is layers 5, 11, 17, 23, 29, 35 and 41, seven in all, and the second clause of the condition (layer_idx >= 42 // 6 × 6 = 42) never fires. The card describes the same thing in words: “5:1 alternating stacking of Kimi Delta Attention (KDA) and MLA”.

  • Ling-3.0-flash Weights availability official claim high confidence
  • Step 3.7 Flash Parameters, licence and weights official claim high confidence

    Parameter count from the safetensors API (BF16 201,365,304,064 plus F32 12,096) rather than the card's rounded 198B. Active parameters are the card's “approximately 11B parameters per token”, stated as an approximation there and carried as one here; config.json gives moe_num_experts 288 and moe_top_k 8 but no active total.

  • Step 3.7 Flash KV-cache geometry estimated high confidence

    2 (key and value) × num_attention_groups 8 × head_dim 128 = 2,048 elements per token per layer, the same on both kinds. The twelve global layers are the whole slope: 49.15 MB per 1,000 tokens at bf16. The thirty-three sliding layers stop at 512 tokens and then hold a fixed 69.21 MB per sequence. Past a few thousand tokens the window is doing almost all of the work — an uncapped reading of this same geometry would be 180.00 KiB per token against the 48.00 KiB per token modelled here. Eight key-value groups means the cache stops dividing beyond eight accelerators and replicates thereafter.

  • Step 3.7 Flash Attention mechanism official claim high confidence

    text_config.layer_types lists 48 entries, 12 of them full_attention at indices 0, 4, 8 ... 44 and the rest sliding_attention. num_hidden_layers is 45 and num_nextn_predict_layers is 3, so indices 45 to 47 are speculative-decoding layers and the model that answers has 12 full and 33 sliding, not 36. Both attention types use num_attention_groups 8 and head_dim 128; attention_other_setting raises the query heads from 64 to 96 on the sliding layers and leaves the key-value geometry alone. sliding_window is 512.

  • Step 3.7 Flash Weights availability official claim high confidence
  • Mistral Small 4 (119B) Parameters, licence and weights official claim high confidence

    119.40B is the safetensors API's dtype split summed (F8_E4M3 117,879,865,344 plus BF16 1,521,452,032), which agrees with the card's “119B parameters, with 6.5B activated per token”. The weights figure follows the same split at one byte and two bytes respectively rather than assuming a single precision.

  • Mistral Small 4 (119B) KV-cache geometry estimated high confidence

    kv_lora_rank 256 plus qk_rope_head_dim 64 = 320 cached elements per token per layer across all 36 layers: 23.04 MB per 1,000 tokens at bf16, or 22.50 KiB per token. Mistral Large 3 caches 576 elements on 61 layers, so the smaller model here is cheaper per token of context by a factor of three from geometry alone. As with every latent model the cache is a single head, so ordinary tensor parallelism replicates it on every rank while the weights divide.

  • Mistral Small 4 (119B) Attention mechanism official claim high confidence

    text_config: kv_lora_rank 256, qk_rope_head_dim 64, q_lora_rank 1024, num_hidden_layers 36, sliding_window null and no layer_types key, so every layer is the same and every layer grows.

  • Mistral Small 4 (119B) Weights availability official claim high confidence
  • DeepSeek V4.1 Flash Parameters, licence and weights official claim high confidence

    TWO PARAMETER COUNTS, AND THIS FIELD HOLDS THE CHECKPOINT ONE. DeepSeek states 552B backbone parameters and 196B Engram parameters as a pair and never sums them; the repository dtype tally for the same files gives 763,205,315,794, which is backbone plus Engram plus the vision tower plus the three DSpark draft layers plus embeddings. The 15.2B that 552 + 196 leaves unaccounted for is not itemised by the publisher and is not guessed here. The checkpoint figure is stored for the reason Qwen3.8-Flash-Next stores 180B rather than its 125B of transformer: a server loads what is in the file. It is also the only one the release is self-consistent with — 510.29 GB over 763.21B is 0.669 bytes per parameter, where a mixed FP4/FP8 release should land, while 552B would put it at 0.924, an outlier against every 4-bit checkpoint in this catalog. ACTIVE PARAMETERS ARE PHASE-DEPENDENT AND THIS FIELD IS NOT. The report states 8B active per prefill token and 16B per decode token, a consequence of the causal encoder-decoder projecting the decoder KV from the encoder final hidden state. 16 is stored because the bandwidth floor this figure feeds is a decode model; the 8B prefill figure is the more distinctive of the two and is recorded here rather than dropped. Artificial Analysis independently lists 552 and 16 for the same model.

  • DeepSeek V4.1 Flash KV-cache geometry estimated high confidence

    Read from inference/model.py in the release repository, and it reproduces the technical report headline figure exactly, which is the check that makes this entry trustworthy rather than plausible. THE LAYER COUNTS HERE ARE BUFFERS, NOT LAYERS THAT ATTEND. config.json lists compress_ratios over 43 layers as 5 at 0, 18 at 2 and 20 at 1, but Attention registers compress_kv_cache only when layer_id is in kv_source_layer_ids [2, 8, 14, 20], and every other layer reads those; the file says so in a comment. So four buffers exist: three at head_dim 512 over ratio 2 = 256 elements per token, and one at 512 over 1 = 512, a 1,280-element slope. THE INDEXER OWNS KEYS ON THE SAME FOUR LAYERS, NOT ON ALL EIGHT index_source_layer_ids: Indexer sets owns_k from kv_source_layers and registers k_cache only then, because the index keys are derived from the compressor latent and only a layer that compresses its own KV can produce them. That is 3 x 128/2 + 1 x 128/1 = 320 elements per token. PRECISION IS PER BUFFER: the latent goes through fp4_act_quant with a block of 16 and an e4m3 scale, so 0.5 bytes plus one scale byte per 16 elements = 0.5625; the index keys use the module-wide fp4_block_size of 32, so 0.53125. 1,280 x 0.5625 + 320 x 0.53125 = 720 + 170 = 890 bytes per token, which is the figure the report states. THE WINDOW IS THE ONLY PART THE PRECISION CONTROL MOVES: window_kv_cache is registered on all 43 layers at [batch, 128, head_dim] and the report says FP8 is retained for it, so it is left on the model default of one byte rather than overridden. 43 x 128 x 512 = 2,818,048 elements, which is 5.64 MB per sequence on the bf16 basis every note in this file quotes for comparability and half of that at the fp8 this profile actually defaults to. RoPE occupies the last 64 of the 512 rather than adding to them, the same slice convention as V4, so an entry is 512 elements and not 576. READING LAYERS ARE NOT BUFFERS EITHER: every layer with a compress ratio above zero reads the buffer of the nearest kv source at or below it (inference/model.py:739-763; vLLM e52be1a6 models/deepseek_v4_1/attention.py:284-289), so 18 layers read the three 2:1 buffers and 20 read the 1:1 one, and each of the eight index_source_layer_ids scores every key its kv source stores (model.py:554-557), so the 2:1 keys are read by 3 layers and layer 20's by 5. Layers 0 and 1 and the three DSpark layers read only their windows. The groups record those counts as readingLayers, which multiply what a step reads and not what memory holds. CORRECTED 2026-09-13: both bandwidth searches charged each buffer once, so the estimate read 890 bytes per token of each copy where a step reads 8,794. STILL EXCLUDED: whether the hierarchical indexer candidate pool (candidate_topk_blocks 2048 at candidate_block_size 8) needs storage beyond the index keys already counted. The release does not say, and the difference between the readings would be larger than the term.

  • DeepSeek V4.1 Flash Attention mechanism official claim high confidence

    The release inference/model.py states the trap in a comment of its own: a compress_ratio above zero does not mean the layer compresses its own KV, only kv_source_layers do and the rest read that same cache. A causal encoder-decoder splits the 40 backbone layers into a 20-layer encoder and a 20-layer decoder, which is why 8B of parameters are active per prefill token and 16B per decode token. WHAT A STEP READS: index_topk counts compressed entries, and the Indexer keeps topk(min(index_topk, end_pos // ratio)) of them (inference/model.py:578), so a ratio-2 layer reads at most 512 entries of two tokens each, which is the 2:1 group's tokensPerEntry, and a ratio-1 layer 512 entries of one token. CORRECTED 2026-09-13: selectedEntries was selectedTokens, and the bandwidth lower bound charged the 2:1 buffers 512 tokens rather than 512 entries.

  • DeepSeek V4.1 Flash Expert routing official claim high confidence

    config.json text_config: n_routed_experts 384, num_experts_per_tok 6, n_shared_experts 1 — the same expert shape as V4 Pro. The three DSpark draft layers run a lighter stack of their own (dspark_n_routed_experts 128, dspark_num_experts_per_tok 3) which is not modelled here.

  • Anthropic · Claude Opus 5 standard · 5m cache write official claim high confidence

    Released 2026-07-24 at the same price as Opus 4.8. Cache write field uses the five-minute tier; the one-hour tier is 2x base input ($10/M) and is not encoded. Anthropic documents the full 1M window at flat pricing, so no long-context surcharge applies. Uses the new-generation tokenizer (~30% more tokens for the same text than Sonnet 4.6 and earlier); raw per-token price comparisons across that boundary undercount effective cost. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included. Anthropic's thinking page lists each model's `max_tokens` ceiling and prints 128k for Claude Opus 5. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Opus 5 standard · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Opus 5 Batch API · 5m cache write official claim high confidence

    Official Batch API 50% multiplier; cache write uses the five-minute tier. New-generation tokenizer (~30% more tokens than Sonnet 4.6). OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included, and 300,000 with the output-300k-2026-03-24 beta header. Anthropic's thinking page prints 128k for Claude Opus 5 and a batches beta ceiling of 300k, and its batch-processing page names Claude Opus 5 among the models for whose batch requests that header “raises the `max_tokens` cap to 300,000”, on the Message Batches API only. An output between the two caps is in doubt rather than refused, because a scenario does not record whether a request sends the header. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward both caps.

  • Anthropic · Claude Opus 5 Batch API · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Opus 4.8 standard · 5m cache write official claim high confidence

    Still offered and supported — Anthropic lists it under “Legacy models (still available)” on its models overview, which is where that heading lives and is now cited rather than the pricing page this row used to point at for it with no deprecation notice — but superseded as the recommended choice by Opus 5 on 2026-07-24 at an identical price, so it is hidden by default. Cache write field uses the five-minute tier. New-generation tokenizer (~30% more tokens than Sonnet 4.6). OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included. Anthropic's thinking page lists each model's `max_tokens` ceiling and prints 128k for Claude Opus 4.8. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Opus 4.8 standard · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Fable 5 standard · 5m cache write official claim high confidence

    Cache write field uses the five-minute tier; one-hour cache writes are not encoded; batch is, in claude-fable-5-batch. Uses the new-generation tokenizer (~30% more tokens for the same text than Sonnet 4.6 and earlier); raw per-token price comparisons across that boundary undercount effective cost. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included. Anthropic's thinking page lists each model's `max_tokens` ceiling and prints 128k for Claude Fable 5. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Fable 5 standard · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Sonnet 4.6 standard · 5m cache write official claim high confidence

    Still offered under “Legacy models (still available)” on Anthropic's models overview with no deprecation notice, but superseded by Sonnet 5 (2026-06-30), so it is hidden by default. Cache write field uses the five-minute tier. Last generation on the older tokenizer. CORRECTED 2026-08-29: this note used to argue that the older tokenizer made its “nominally identical $3/$15 cheaper in practice than the Sonnet 5 standard rate”. Both halves were wrong and together they reversed the conclusion. Sonnet 5's standard rate is $2/$10, in this catalog's own claude-sonnet-5 row and on Anthropic's page; the only $3/$15 is claude-sonnet-5-standard, which this catalog itself describes as a price that never took effect. And the tokenizer difference is about 23%, not 30%: the newer tokenizer produces about 30% more tokens for the same text, so the older one produces 1/1.30 of them. $3 over 0.77 as many tokens is an effective $2.31 against Sonnet 5's $2.00, so Sonnet 4.6 is about 15% dearer for the same text, not cheaper. The heading quoted above was also cited to the pricing page, which does not carry it. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included. Anthropic's thinking page lists each model's `max_tokens` ceiling and prints 128k for Claude Sonnet 4.6. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Sonnet 4.6 standard · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Opus 4.8 Batch API · 5m cache write official claim high confidence

    Official Batch API 50% multiplier; cache write uses the five-minute tier. Uses the new-generation tokenizer (~30% more tokens for the same text than Sonnet 4.6 and earlier); raw per-token price comparisons across that boundary undercount effective cost. Retired 2026-08-15: Claude Opus 5 Batch charges exactly the same /5 and scores 63 against 57 on the same index, so there is no workload where this tier is the better buy. The standard 4.8 tier was retired at the Opus 5 launch and this one was left behind — a model is not half-retired. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included, and 300,000 with the output-300k-2026-03-24 beta header. Anthropic's thinking page prints 128k for Claude Opus 4.8 and a batches beta ceiling of 300k, and its batch-processing page names Claude Opus 4.8 among the models for whose batch requests that header “raises the `max_tokens` cap to 300,000”, on the Message Batches API only. An output between the two caps is in doubt rather than refused, because a scenario does not record whether a request sends the header. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward both caps.

  • Anthropic · Claude Opus 4.8 Batch API · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Fable 5 Batch API · 5m cache write official claim high confidence

    Official Batch API 50% multiplier; cache write uses the five-minute tier. Uses the new-generation tokenizer (~30% more tokens for the same text than Sonnet 4.6 and earlier); raw per-token price comparisons across that boundary undercount effective cost. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included, and no higher cap on the Batch API. Anthropic's thinking page prints 128k for Claude Fable 5 and a dash in its batches beta ceiling column, and its batch-processing page does not name Claude Fable 5 among the models the output-300k-2026-03-24 header raises. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Fable 5 Batch API · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Sonnet 4.6 Batch API · 5m cache write official claim high confidence

    Official Batch API 50% multiplier; cache write uses the five-minute tier. Superseded by the Sonnet 5 batch tier. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included, and 300,000 with the output-300k-2026-03-24 beta header. Anthropic's thinking page prints 128k for Claude Sonnet 4.6 and a batches beta ceiling of 300k, and its batch-processing page names Claude Sonnet 4.6 among the models for whose batch requests that header “raises the `max_tokens` cap to 300,000”, on the Message Batches API only. An output between the two caps is in doubt rather than refused, because a scenario does not record whether a request sends the header. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward both caps.

  • Anthropic · Claude Sonnet 4.6 Batch API · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Sonnet 5 standard · 5m cache write official claim high confidence

    CHANGED at the 2026-08-13 refresh: this is now the permanent standard rate. Anthropic: "The $2/$10 per million input/output token pricing for Claude Sonnet 5, announced at launch as introductory pricing through August 31, 2026, is now the standard price. The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur." The expiry this catalog previously carried has therefore been removed. Five-minute cache-write tier (one-hour is $4/M). New-generation tokenizer (~30% more tokens than Sonnet 4.6). OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included. Anthropic's thinking page lists each model's `max_tokens` ceiling and prints 128k for Claude Sonnet 5. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Sonnet 5 standard · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Sonnet 5 cancelled 2026-09-01 increase official claim high confidence

    A price that never took effect. Anthropic scheduled Sonnet 5 to rise to $3/$15 on 2026-09-01, this catalog encoded it so multi-year scenarios could be priced honestly against it, and Anthropic then cancelled the increase. Kept only so scenario URLs shared while the increase was still scheduled continue to resolve; it is hidden by default and should not be used for new planning. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included. Anthropic's thinking page lists each model's `max_tokens` ceiling and prints 128k for Claude Sonnet 5. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Sonnet 5 cancelled 2026-09-01 increase · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Sonnet 5 Batch API · 5m cache write official claim high confidence

    Now the permanent standard rate with the official Batch API 50% multiplier — Anthropic cancelled the 2026-09-01 increase that this tier's expiry previously anticipated. Five-minute cache-write tier. New-generation tokenizer (~30% more tokens than Sonnet 4.6). OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included, and 300,000 with the output-300k-2026-03-24 beta header. Anthropic's thinking page prints 128k for Claude Sonnet 5 and a batches beta ceiling of 300k, and its batch-processing page names Claude Sonnet 5 among the models for whose batch requests that header “raises the `max_tokens` cap to 300,000”, on the Message Batches API only. An output between the two caps is in doubt rather than refused, because a scenario does not record whether a request sends the header. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward both caps.

  • Anthropic · Claude Sonnet 5 Batch API · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Sonnet 5 Batch · cancelled 2026-09-01 increase official claim high confidence

    The batch half of a price increase Anthropic scheduled for 2026-09-01 and then cancelled. Retained only so older shared scenarios resolve. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included, and 300,000 with the output-300k-2026-03-24 beta header. Anthropic's thinking page prints 128k for Claude Sonnet 5 and a batches beta ceiling of 300k, and its batch-processing page names Claude Sonnet 5 among the models for whose batch requests that header “raises the `max_tokens` cap to 300,000”, on the Message Batches API only. An output between the two caps is in doubt rather than refused, because a scenario does not record whether a request sends the header. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward both caps.

  • Anthropic · Claude Sonnet 5 Batch · cancelled 2026-09-01 increase · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Haiku 4.5 standard · 5m cache write official claim high confidence

    OUTPUT CAP, recorded 2026-09-13: 64,000 tokens, reasoning included. Anthropic's thinking page lists each model's `max_tokens` ceiling and prints 64k for Claude Haiku 4.5, with its batches beta ceiling “Not available”. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Haiku 4.5 standard · 5m cache write · data residency official claim high confidence

    Claude Haiku 4.5 predates the parameter. "Requests with inference_geo on Claude Opus 4.5, Claude Sonnet 4.5, Claude Haiku 4.5, or earlier models return a 400 error" — so this tier cannot be pinned to a jurisdiction at any price, which is a different answer from the surcharge being unknown. Recorded so the calculator can say which. Read 2026-09-10.

  • OpenAI · GPT-5.6 Sol standard official claim high confidence

    CUT 2026-08-21 to $4.00 / $0.40 / $5.00 / $20.00 from $5.00 / $0.50 / $6.25 / $30.00 — 20% off input and 33.3% off output. OpenAI's API changelog, entry dated Aug 21 and tagged gpt-5.6-sol: "GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens, representing 20% lower input pricing and 33% lower output pricing." The cache write stays at 1.25x input, which is what the reversal below established. Deliberately carries no effectiveUntil. OpenAI's own wording is "GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026" — a floor on how long the rate lasts, not a date it stops. Writing 2026-11-21 into effectiveUntil would assert an increase the vendor has not announced. CORRECTED 2026-08-27: this note also claimed the picker would then hide a live tier, which is false about this repository's own code — the picker filters on whether a successor row exists (listPriceTierFor), never on a date, so a tier with an end date and no successor stays visible. The real cost of leaving the field empty is the opposite and worth stating: no promotional-pricing caution fires, so a multi-year break-even against this tier is computed against a rate the vendor guarantees only to 2026-11-21 and says nothing about afterwards. Contrast GLM-5.3-Flash, where Z.ai wrote "the promotion ends at 24:00 on September 9, 2026" and effectiveUntil is the right field. Terra and Luna were re-read from the same table on 2026-08-27 and are unchanged at $2.00/$12.00 and $0.20/$1.20. Explicitly excluded from the 2026-07-30 price cut that reduced Terra 20% and Luna 80%. The >272K long-context multipliers apply to input, cached input, cache write and output alike. CORRECTED 2026-08-15: this tier carried a cache-write rate of $6.25/M, 1.25x its input rate, copied from Anthropic’s billing model. OpenAI does not bill a cache write at all — caching is automatic, the first call pays the ordinary input rate and matched prefixes pay the cached rate afterwards. The write bucket is therefore priced at the plain input rate, which is what actually happens. Corroborated by the owner’s own Codex telemetry, where the cache-write counter is structurally zero across 357 sessions and 6.12 billion input tokens. REVERSED 2026-08-21, and this is the important half of the note. The "correction" above was wrong: OpenAI does bill a cache write on this model family, at exactly the 1.25x multiplier it was accused of borrowing from Anthropic. Their caching guide states it outright. Under an expander headed “GPT-5.6 and later”: “For GPT-5.6 and later, cache writes cost 1.25× the standard, uncached input-token rate.” Its comparison table says the same in a Cache write charge row, reading “1.25× the uncached input-token rate” for that family against “No additional cache-write charge” for the two older ones. CORRECTED 2026-08-27: the sentence quoted here until today was a paraphrase in quotation marks — none of “have no additional fee”, “billed at 1.25” or “reported in cache_write_tokens” appears anywhere on that page. It said the right thing and it was not a quotation, which on a site whose premise is a citable audit trail is its own defect. The same correction removes the absolute figures that followed it — $6.25, $2.50 and $0.25 against inputs of $5.00, $2.00 and $0.20 — because that was the table as it read on 2026-08-21 and Sol was cut on the same day. The ratio is what survived. The telemetry that seemed to corroborate the error is not wrong either; it is pre-GPT-5.6 traffic, on models where the fee genuinely did not exist, and it was generalised to a family that had introduced one on 2026-07-09. A measurement of the wrong thing is more persuasive than no measurement at all, which is exactly why it took a source read to catch. TOKEN LIMITS, recorded 2026-09-13: a maximum input of 922,000 tokens and an output cap of 128,000 inside the 1,050,000 window, from the model details on OpenAI's gpt-5.6-sol page: “Maximum input tokens: 922,000” and “128,000 max output tokens”. Reasoning counts toward the output cap: OpenAI's token-counting guide says the `max_output_tokens` and `max_completion_tokens` parameters “limit all tokens generated by the model, including non-visible tokens.”

  • OpenAI · GPT-5.6 Sol standard · data residency official claim high confidence

    OpenAI states: "Data residency endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency." The support table on the same page lists this model as eligible, and every OpenAI tier in this catalog is priced from 2026-07-30 or later, so both conditions hold. Two things OpenAI does not say, and this catalog does not invent: which token categories the uplift covers - Anthropic enumerates input, output, cache reads and cache writes, OpenAI writes only "10% uplift" against the endpoint - and how it composes with the Batch and Flex discounts. Read 2026-09-10.

  • OpenAI · GPT-5.6 Terra standard official claim high confidence

    Cut 20% on 2026-07-30 from $2.50/$15.00. OpenAI: "Starting July 30, API pricing is $2 per million input tokens and $12 per million output tokens for Terra." The >272K long-context multipliers apply to input, cached input, cache write and output alike. CORRECTED 2026-08-15: this tier carried a cache-write rate of $2.5/M, 1.25x its input rate, copied from Anthropic’s billing model. OpenAI does not bill a cache write at all — caching is automatic, the first call pays the ordinary input rate and matched prefixes pay the cached rate afterwards. The write bucket is therefore priced at the plain input rate, which is what actually happens. Corroborated by the owner’s own Codex telemetry, where the cache-write counter is structurally zero across 357 sessions and 6.12 billion input tokens. REVERSED 2026-08-21, and this is the important half of the note. The "correction" above was wrong: OpenAI does bill a cache write on this model family, at exactly the 1.25x multiplier it was accused of borrowing from Anthropic. Their caching guide states it outright. Under an expander headed “GPT-5.6 and later”: “For GPT-5.6 and later, cache writes cost 1.25× the standard, uncached input-token rate.” Its comparison table says the same in a Cache write charge row, reading “1.25× the uncached input-token rate” for that family against “No additional cache-write charge” for the two older ones. CORRECTED 2026-08-27: the sentence quoted here until today was a paraphrase in quotation marks — none of “have no additional fee”, “billed at 1.25” or “reported in cache_write_tokens” appears anywhere on that page. It said the right thing and it was not a quotation, which on a site whose premise is a citable audit trail is its own defect. The same correction removes the absolute figures that followed it — $6.25, $2.50 and $0.25 against inputs of $5.00, $2.00 and $0.20 — because that was the table as it read on 2026-08-21 and Sol was cut on the same day. The ratio is what survived. The telemetry that seemed to corroborate the error is not wrong either; it is pre-GPT-5.6 traffic, on models where the fee genuinely did not exist, and it was generalised to a family that had introduced one on 2026-07-09. A measurement of the wrong thing is more persuasive than no measurement at all, which is exactly why it took a source read to catch. TOKEN LIMITS, recorded 2026-09-13: a maximum input of 922,000 tokens and an output cap of 128,000 inside the 1,050,000 window, from the model details on OpenAI's gpt-5.6-terra page: “Maximum input tokens: 922,000” and “128,000 max output tokens”. Reasoning counts toward the output cap: OpenAI's token-counting guide says the `max_output_tokens` and `max_completion_tokens` parameters “limit all tokens generated by the model, including non-visible tokens.”

  • OpenAI · GPT-5.6 Terra standard · data residency official claim high confidence

    OpenAI states: "Data residency endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency." The support table on the same page lists this model as eligible, and every OpenAI tier in this catalog is priced from 2026-07-30 or later, so both conditions hold. Two things OpenAI does not say, and this catalog does not invent: which token categories the uplift covers - Anthropic enumerates input, output, cache reads and cache writes, OpenAI writes only "10% uplift" against the endpoint - and how it composes with the Batch and Flex discounts. Read 2026-09-10.

  • OpenAI · GPT-5.6 Luna standard official claim high confidence

    Cut 80% on 2026-07-30 from $1.00/$6.00. OpenAI: "GPT-5.6 Luna, our fastest and most affordable model, will cost 80% less." At $0.20/$1.20 this is the cheapest frontier-lab tier in the catalog and the hardest comparator for any self-hosting case to beat. CORRECTED 2026-08-15: this tier carried a cache-write rate of $0.25/M, 1.25x its input rate, copied from Anthropic’s billing model. OpenAI does not bill a cache write at all — caching is automatic, the first call pays the ordinary input rate and matched prefixes pay the cached rate afterwards. The write bucket is therefore priced at the plain input rate, which is what actually happens. Corroborated by the owner’s own Codex telemetry, where the cache-write counter is structurally zero across 357 sessions and 6.12 billion input tokens. REVERSED 2026-08-21, and this is the important half of the note. The "correction" above was wrong: OpenAI does bill a cache write on this model family, at exactly the 1.25x multiplier it was accused of borrowing from Anthropic. Their caching guide states it outright. Under an expander headed “GPT-5.6 and later”: “For GPT-5.6 and later, cache writes cost 1.25× the standard, uncached input-token rate.” Its comparison table says the same in a Cache write charge row, reading “1.25× the uncached input-token rate” for that family against “No additional cache-write charge” for the two older ones. CORRECTED 2026-08-27: the sentence quoted here until today was a paraphrase in quotation marks — none of “have no additional fee”, “billed at 1.25” or “reported in cache_write_tokens” appears anywhere on that page. It said the right thing and it was not a quotation, which on a site whose premise is a citable audit trail is its own defect. The same correction removes the absolute figures that followed it — $6.25, $2.50 and $0.25 against inputs of $5.00, $2.00 and $0.20 — because that was the table as it read on 2026-08-21 and Sol was cut on the same day. The ratio is what survived. The telemetry that seemed to corroborate the error is not wrong either; it is pre-GPT-5.6 traffic, on models where the fee genuinely did not exist, and it was generalised to a family that had introduced one on 2026-07-09. A measurement of the wrong thing is more persuasive than no measurement at all, which is exactly why it took a source read to catch. TOKEN LIMITS, recorded 2026-09-13: a maximum input of 922,000 tokens and an output cap of 128,000 inside the 1,050,000 window, from the model details on OpenAI's gpt-5.6-luna page: “Maximum input tokens: 922,000” and “128,000 max output tokens”. Reasoning counts toward the output cap: OpenAI's token-counting guide says the `max_output_tokens` and `max_completion_tokens` parameters “limit all tokens generated by the model, including non-visible tokens.”

  • OpenAI · GPT-5.6 Luna standard · data residency official claim high confidence

    OpenAI states: "Data residency endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency." The support table on the same page lists this model as eligible, and every OpenAI tier in this catalog is priced from 2026-07-30 or later, so both conditions hold. Two things OpenAI does not say, and this catalog does not invent: which token categories the uplift covers - Anthropic enumerates input, output, cache reads and cache writes, OpenAI writes only "10% uplift" against the endpoint - and how it composes with the Batch and Flex discounts. Read 2026-09-10.

  • OpenAI · GPT-5.6 Sol Flex / Batch · 50% multiplier official claim high confidence

    CUT 2026-08-21 to $4.00 / $0.40 / $5.00 / $20.00 from $5.00 / $0.50 / $6.25 / $30.00 — 20% off input and 33.3% off output. OpenAI's API changelog, entry dated Aug 21 and tagged gpt-5.6-sol: "GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens, representing 20% lower input pricing and 33% lower output pricing." The cache write stays at 1.25x input, which is what the reversal below established. Deliberately carries no effectiveUntil. OpenAI's own wording is "GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026" — a floor on how long the rate lasts, not a date it stops. Writing 2026-11-21 into effectiveUntil would assert an increase the vendor has not announced. CORRECTED 2026-08-27: this note also claimed the picker would then hide a live tier, which is false about this repository's own code — the picker filters on whether a successor row exists (listPriceTierFor), never on a date, so a tier with an end date and no successor stays visible. The real cost of leaving the field empty is the opposite and worth stating: no promotional-pricing caution fires, so a multi-year break-even against this tier is computed against a rate the vendor guarantees only to 2026-11-21 and says nothing about afterwards. Contrast GLM-5.3-Flash, where Z.ai wrote "the promotion ends at 24:00 on September 9, 2026" and effectiveUntil is the right field. Terra and Luna were re-read from the same table on 2026-08-27 and are unchanged at $2.00/$12.00 and $0.20/$1.20. Sol was explicitly excluded from the 2026-07-30 cut ("Sol pricing remains unchanged"). OpenAI documents Flex tokens as priced at Batch API rates and the Batch discount as a flat 50%, so both tiers share this row. CORRECTED 2026-08-15: this tier carried a cache-write rate of $6.25/M, 1.25x its input rate, copied from Anthropic’s billing model. OpenAI does not bill a cache write at all — caching is automatic, the first call pays the ordinary input rate and matched prefixes pay the cached rate afterwards. The write bucket is therefore priced at the plain input rate, which is what actually happens. Corroborated by the owner’s own Codex telemetry, where the cache-write counter is structurally zero across 357 sessions and 6.12 billion input tokens. REVERSED 2026-08-21, and this is the important half of the note. The "correction" above was wrong: OpenAI does bill a cache write on this model family, at exactly the 1.25x multiplier it was accused of borrowing from Anthropic. Their caching guide states it outright. Under an expander headed “GPT-5.6 and later”: “For GPT-5.6 and later, cache writes cost 1.25× the standard, uncached input-token rate.” Its comparison table says the same in a Cache write charge row, reading “1.25× the uncached input-token rate” for that family against “No additional cache-write charge” for the two older ones. CORRECTED 2026-08-27: the sentence quoted here until today was a paraphrase in quotation marks — none of “have no additional fee”, “billed at 1.25” or “reported in cache_write_tokens” appears anywhere on that page. It said the right thing and it was not a quotation, which on a site whose premise is a citable audit trail is its own defect. The same correction removes the absolute figures that followed it — $6.25, $2.50 and $0.25 against inputs of $5.00, $2.00 and $0.20 — because that was the table as it read on 2026-08-21 and Sol was cut on the same day. The ratio is what survived. The telemetry that seemed to corroborate the error is not wrong either; it is pre-GPT-5.6 traffic, on models where the fee genuinely did not exist, and it was generalised to a family that had introduced one on 2026-07-09. A measurement of the wrong thing is more persuasive than no measurement at all, which is exactly why it took a source read to catch. TOKEN LIMITS, recorded 2026-09-13: a maximum input of 922,000 tokens and an output cap of 128,000 inside the 1,050,000 window, from the model details on OpenAI's gpt-5.6-sol page: “Maximum input tokens: 922,000” and “128,000 max output tokens”. Reasoning counts toward the output cap: OpenAI's token-counting guide says the `max_output_tokens` and `max_completion_tokens` parameters “limit all tokens generated by the model, including non-visible tokens.”

  • OpenAI · GPT-5.6 Sol Flex / Batch · 50% multiplier · data residency estimated medium confidence

    OpenAI states: "Data residency endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency." The support table on the same page lists this model as eligible, and every OpenAI tier in this catalog is priced from 2026-07-30 or later, so both conditions hold. Two things OpenAI does not say, and this catalog does not invent: which token categories the uplift covers - Anthropic enumerates input, output, cache reads and cache writes, OpenAI writes only "10% uplift" against the endpoint - and how it composes with the Batch and Flex discounts. Read 2026-09-10. Applied multiplicatively on top of this arrangement discount, which is a derivation: OpenAI states the uplift and the discount separately and never their interaction. Anthropic documents multiplicative composition for the equivalent case, which is why it is the reading taken, and is not a statement about OpenAI.

  • OpenAI · GPT-5.6 Terra Flex / Batch · 50% multiplier official claim high confidence

    Reflects the 2026-07-30 cut. OpenAI documents Flex tokens as priced at Batch API rates, and the Batch discount as a flat 50%, so both tiers share this row. Long-context multipliers still apply to the full request. CORRECTED 2026-08-15: this tier carried a cache-write rate of $2.5/M, 1.25x its input rate, copied from Anthropic’s billing model. OpenAI does not bill a cache write at all — caching is automatic, the first call pays the ordinary input rate and matched prefixes pay the cached rate afterwards. The write bucket is therefore priced at the plain input rate, which is what actually happens. Corroborated by the owner’s own Codex telemetry, where the cache-write counter is structurally zero across 357 sessions and 6.12 billion input tokens. REVERSED 2026-08-21, and this is the important half of the note. The "correction" above was wrong: OpenAI does bill a cache write on this model family, at exactly the 1.25x multiplier it was accused of borrowing from Anthropic. Their caching guide states it outright. Under an expander headed “GPT-5.6 and later”: “For GPT-5.6 and later, cache writes cost 1.25× the standard, uncached input-token rate.” Its comparison table says the same in a Cache write charge row, reading “1.25× the uncached input-token rate” for that family against “No additional cache-write charge” for the two older ones. CORRECTED 2026-08-27: the sentence quoted here until today was a paraphrase in quotation marks — none of “have no additional fee”, “billed at 1.25” or “reported in cache_write_tokens” appears anywhere on that page. It said the right thing and it was not a quotation, which on a site whose premise is a citable audit trail is its own defect. The same correction removes the absolute figures that followed it — $6.25, $2.50 and $0.25 against inputs of $5.00, $2.00 and $0.20 — because that was the table as it read on 2026-08-21 and Sol was cut on the same day. The ratio is what survived. The telemetry that seemed to corroborate the error is not wrong either; it is pre-GPT-5.6 traffic, on models where the fee genuinely did not exist, and it was generalised to a family that had introduced one on 2026-07-09. A measurement of the wrong thing is more persuasive than no measurement at all, which is exactly why it took a source read to catch. TOKEN LIMITS, recorded 2026-09-13: a maximum input of 922,000 tokens and an output cap of 128,000 inside the 1,050,000 window, from the model details on OpenAI's gpt-5.6-terra page: “Maximum input tokens: 922,000” and “128,000 max output tokens”. Reasoning counts toward the output cap: OpenAI's token-counting guide says the `max_output_tokens` and `max_completion_tokens` parameters “limit all tokens generated by the model, including non-visible tokens.”

  • OpenAI · GPT-5.6 Terra Flex / Batch · 50% multiplier · data residency estimated medium confidence

    OpenAI states: "Data residency endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency." The support table on the same page lists this model as eligible, and every OpenAI tier in this catalog is priced from 2026-07-30 or later, so both conditions hold. Two things OpenAI does not say, and this catalog does not invent: which token categories the uplift covers - Anthropic enumerates input, output, cache reads and cache writes, OpenAI writes only "10% uplift" against the endpoint - and how it composes with the Batch and Flex discounts. Read 2026-09-10. Applied multiplicatively on top of this arrangement discount, which is a derivation: OpenAI states the uplift and the discount separately and never their interaction. Anthropic documents multiplicative composition for the equivalent case, which is why it is the reading taken, and is not a statement about OpenAI.

  • OpenAI · GPT-5.6 Luna Flex / Batch · 50% multiplier official claim high confidence

    Reflects the 2026-07-30 cut. Flex and Batch share the published flat 50% multiplier. At $0.10/$0.60 effective, this is the cheapest tier in the catalog by a wide margin. CORRECTED 2026-08-15: this tier carried a cache-write rate of $0.25/M, 1.25x its input rate, copied from Anthropic’s billing model. OpenAI does not bill a cache write at all — caching is automatic, the first call pays the ordinary input rate and matched prefixes pay the cached rate afterwards. The write bucket is therefore priced at the plain input rate, which is what actually happens. Corroborated by the owner’s own Codex telemetry, where the cache-write counter is structurally zero across 357 sessions and 6.12 billion input tokens. REVERSED 2026-08-21, and this is the important half of the note. The "correction" above was wrong: OpenAI does bill a cache write on this model family, at exactly the 1.25x multiplier it was accused of borrowing from Anthropic. Their caching guide states it outright. Under an expander headed “GPT-5.6 and later”: “For GPT-5.6 and later, cache writes cost 1.25× the standard, uncached input-token rate.” Its comparison table says the same in a Cache write charge row, reading “1.25× the uncached input-token rate” for that family against “No additional cache-write charge” for the two older ones. CORRECTED 2026-08-27: the sentence quoted here until today was a paraphrase in quotation marks — none of “have no additional fee”, “billed at 1.25” or “reported in cache_write_tokens” appears anywhere on that page. It said the right thing and it was not a quotation, which on a site whose premise is a citable audit trail is its own defect. The same correction removes the absolute figures that followed it — $6.25, $2.50 and $0.25 against inputs of $5.00, $2.00 and $0.20 — because that was the table as it read on 2026-08-21 and Sol was cut on the same day. The ratio is what survived. The telemetry that seemed to corroborate the error is not wrong either; it is pre-GPT-5.6 traffic, on models where the fee genuinely did not exist, and it was generalised to a family that had introduced one on 2026-07-09. A measurement of the wrong thing is more persuasive than no measurement at all, which is exactly why it took a source read to catch. TOKEN LIMITS, recorded 2026-09-13: a maximum input of 922,000 tokens and an output cap of 128,000 inside the 1,050,000 window, from the model details on OpenAI's gpt-5.6-luna page: “Maximum input tokens: 922,000” and “128,000 max output tokens”. Reasoning counts toward the output cap: OpenAI's token-counting guide says the `max_output_tokens` and `max_completion_tokens` parameters “limit all tokens generated by the model, including non-visible tokens.”

  • OpenAI · GPT-5.6 Luna Flex / Batch · 50% multiplier · data residency estimated medium confidence

    OpenAI states: "Data residency endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency." The support table on the same page lists this model as eligible, and every OpenAI tier in this catalog is priced from 2026-07-30 or later, so both conditions hold. Two things OpenAI does not say, and this catalog does not invent: which token categories the uplift covers - Anthropic enumerates input, output, cache reads and cache writes, OpenAI writes only "10% uplift" against the endpoint - and how it composes with the Batch and Flex discounts. Read 2026-09-10. Applied multiplicatively on top of this arrangement discount, which is a derivation: OpenAI states the uplift and the discount separately and never their interaction. Anthropic documents multiplicative composition for the equivalent case, which is why it is the reading taken, and is not a statement about OpenAI.

  • Google · Gemini 2.5 Pro ≤200K official claim medium confidence

    CORRECTED at the 2026-08-14 refresh: the 2026-10-16 shutdown date this catalog carried is no longer what Google publishes. Its deprecations page now shows the stable gemini-2.5-pro as "No shutdown date announced", so the date is removed rather than left standing as a deadline nobody has committed to. The entry stays superseded because a newer generation exists and Google points migrations at gemini-3.1-pro-preview — which is itself still Preview, so there is still no GA Pro tier in the Gemini 3 family to move to. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-2.5-pro. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Google · Gemini 2.5 Pro Batch · ≤200K official claim medium confidence

    Official Batch API and context-caching prices; above 200K the configured 2x input multiplier also yields the published $0.25/M cached rate. The batch discount does not halve the cached-input rate. Superseded because a newer generation exists; the 2026-10-16 shutdown date previously recorded here was removed at the 2026-08-14 refresh after Google's deprecations page reverted to "No shutdown date announced". TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-2.5-pro. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Google · Gemini 3.1 Pro Preview · ≤200K official claim medium confidence

    Preview (non-GA) tier, and as of 2026-08-01 still the only Gemini 3 Pro tier that exists — Google's own migration path off the retiring 2.5 Pro points here. Above 200K: $4 in / $18 out / $0.40 cached, matching the configured multipliers. Context window confirmed at 1,048,576 input tokens (65,536 output) at the 2026-08-01 refresh, having been omitted rather than guessed previously. Context caching adds a $4.50 per million tokens per hour storage fee that this calculator does not model. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-3.1-pro-preview. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Google · Gemini 3.6 Flash standard official claim high confidence

    CORRECTED TWICE. The 2026-08-13 refresh caught that this tier was carried at Google's 2027 list rate, double what is charged, but only the cached figure was actually updated — input and output stayed at $1.50/$7.50 while the note claimed otherwise. Both are now $0.75/$3.75, matching the live page's "through December 31, 2026" promotional rate. Superseded by Gemini 3.7 Flash, which Google released 2026-08-13 at an identical price and which wins every published benchmark against this tier (DeepSWE 65.3% vs 48.6%, FrontierCode 43.6% vs 34.4%, GDM-MRCR 97.0% vs 91.8%) — there is no workload where paying the same for this one is the better choice. Google has published no shutdown date for it. No long-context tiering. Context-cache storage of $1 per million tokens per hour is not modeled. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-3.6-flash. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Google · Gemini 3.6 Flash list rate · from 2027-01-01 official claim high confidence

    The rate Google has published as taking effect 2027-01-01, carried separately so a multi-year scenario can be priced against what will actually be charged rather than against a promotion that ends inside the horizon. Exactly double the current promotional rate on every category. Retired 2026-08-15 alongside the promotional 3.6 tier it belongs to: the 3.7 standing rate is identical at $1.50/$7.50 and scores 56 against 52. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-3.6-flash. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Google · Gemini 3.7 Flash standard · promotional through 2026-12-31 official claim high confidence

    Released GA on 2026-08-13 and now Google's current Flash tier. Priced identically to 3.6 Flash — $0.75/$3.75/$0.075 through 2026-12-31, doubling on 2027-01-01 — while winning every published benchmark against it, so 3.6 is marked superseded. Token limits of 1,048,576 input and 65,536 output, from Google's API model reference, gemini-3-7-flash-doc; the pricing page states neither. CORRECTED 2026-09-13: this sentence credited the 1,048,576 to the DeepMind model card, which says only “a token context window of up to 1M”, beside “a 64K token output”. The card's own source record was corrected on 2026-08-29 and this sentence was not. TOKEN LIMITS, recorded 2026-09-13: maxInputTokens and maxOutputTokens hold the model reference's two limits. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.” No long-context tiering. OMISSION, recorded 2026-08-21: Google bills no cache write, which is why that field is absent — but it does bill cache *storage* by the hour, at $0.50 per 1M tokens rising to $1.00 on 2027-01-01. Nothing in this calculator models a per-hour storage fee, so a workload holding a large prefix warm across idle time is priced low here. The Gemini 3.1 Pro entry already disclosed its own equivalent; this one did not.

  • Google · Gemini 3.7 Flash Batch / Flex · 50% multiplier · promotional through 2026-12-31 official claim high confidence

    Google publishes Batch and Flex rows for this model at exactly half the standard rate on all three categories — input, output and context caching alike — so a single 0.5 multiplier is the right shape here. It is not the right shape for every Google model: Gemini 2.5 Pro batch halves input and output but charges the same context-caching rate as its standard tier, which is why that entry carries explicit per-category rates instead. Cache storage is never discounted on any of them. ADDED 2026-08-21. A batch arrangement was published for the catalog's headline Google tier and was not carried, so a reader comparing deferrable work against this model was being shown twice the rate Google charges for it. Flex is priced identically to Batch and shares this row, the same way OpenAI's two are shared. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-3.7-flash. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Google · Gemini 3.7 Flash list rate · from 2027-01-01 official claim high confidence

    The rate Google publishes as taking effect 2027-01-01, exactly double the current promotional rate on every category. Carried separately so a multi-year horizon can be priced against what will be charged rather than a promotion that ends inside it — and so the "price at standing rates" switch has something to switch to. Whether the increase actually happens is a forecast: Anthropic scheduled a comparable rise for Sonnet 5 and then cancelled it. OMISSION, recorded 2026-08-21: Google bills no cache write, which is why that field is absent — but it does bill cache *storage* by the hour, at $0.50 per 1M tokens rising to $1.00 on 2027-01-01. Nothing in this calculator models a per-hour storage fee, so a workload holding a large prefix warm across idle time is priced low here. The Gemini 3.1 Pro entry already disclosed its own equivalent; this one did not. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-3.7-flash. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Google · Gemini 3.7 Flash Batch / Flex · 50% multiplier · list rate from 2027-01-01 official claim high confidence

    Google publishes Batch and Flex rows for this model at exactly half the standard rate on all three categories — input, output and context caching alike — so a single 0.5 multiplier is the right shape here. It is not the right shape for every Google model: Gemini 2.5 Pro batch halves input and output but charges the same context-caching rate as its standard tier, which is why that entry carries explicit per-category rates instead. Cache storage is never discounted on any of them. ADDED 2026-08-21. A batch arrangement was published for the catalog's headline Google tier and was not carried, so a reader comparing deferrable work against this model was being shown twice the rate Google charges for it. Flex is priced identically to Batch and shares this row, the same way OpenAI's two are shared. This row carries the post-2027 list rate, so the "price at standing rates" switch has somewhere to send a batch scenario as well as a standard one. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-3.7-flash. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Moonshot AI · Kimi K3 standard official claim high confidence

    Hit/miss pricing encoded as cached input and uncached input; the hit/miss model has no separate cache-write fee, so writes are priced at standard input. Reconfirmed 2026-07-24. CORRECTED 2026-08-27: this tier carried 1,000,000 tokens of context while the local kimi-k3 entry for the same model carried 1,048,576, so one page of the site contradicted another. Moonshot's own pricing table settles it: the row for kimi-k3 has a Context Window column whose cell reads "1,048,576 tokens" in full. The "1M" that appears in the prose beside it is a pricing unit, not the window — the same page says "Here, 1M = 1,000,000. The prices in the table represent the cost per 1M tokens consumed." Read off the raw HTML rather than a rendered summary, because the figure lives inside the page's own JSX payload. Note this is the opposite outcome from Command A+, where the vendor writes a decimal 200000 in its own config and the decimal reading is therefore the right one: what settles the question is which number the vendor itself spells out, not which is rounder. OUTPUT LIMIT, read 2026-09-13: none is published below the window. Moonshot's Kimi K3 quickstart says “`max_completion_tokens` defaults to 131072 and can be set up to 1048576.” That ceiling is the window itself, so the window is the only limit recorded, and 131,072 is a default a request can raise rather than a cap.

  • xAI · Grok 4.6 standard · auto long-context ≥200K official claim high confidence

    Released 2026-08-12 on the same 1.5T foundation as Grok 4.5 — xAI attributes the gains to longer post-training, not a larger model. Input and output are unchanged from 4.5, but cached input rose from $0.30 to $0.50, so a cache-heavy workload is the one case where the newer model costs more. Uniform 2x above 200K on input, cached input and output. No batch discount is published for this tier (xAI lists one only for the Grok 4.3 and 4.20 lines). CORRECTED 2026-08-21: the long-context boundary is inclusive here and the engine treated it as strict, so a prompt of exactly 200,000 tokens was billed at the single rate. xAI: "requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request", which is also what this tier's own label says. Also recorded, because the note previously said only that no batch discount is published: xAI states that "grok-4.6 and grok-4.5 are not currently supported for Batch API requests and will be rejected". Batch is unavailable here, not merely undiscounted. Where xAI does offer it the discount is 20% rather than 50%, applying to input, output, cached and reasoning tokens alike. OUTPUT LIMIT, read 2026-09-13: xAI's Grok 4.6 guide gives the output limit as “No text output limit” and the context window as 500,000 tokens, so the window is the only limit recorded.

  • xAI · Grok 4.5 standard · auto long-context ≥200K official claim high confidence

    Superseded as flagship by Grok 4.6 on 2026-08-12, but still fully listed and priced by xAI with no deprecation tag — and its $0.30 cached input is cheaper than its successor's $0.50. Uniform 2x above the 200K threshold. Note the 500K window is smaller than the older Grok 4.3 line's 1M. CORRECTED 2026-08-21: the long-context boundary is inclusive here and the engine treated it as strict, so a prompt of exactly 200,000 tokens was billed at the single rate. xAI: "requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request", which is also what this tier's own label says. Also recorded, because the note previously said only that no batch discount is published: xAI states that "grok-4.6 and grok-4.5 are not currently supported for Batch API requests and will be rejected". Batch is unavailable here, not merely undiscounted. Where xAI does offer it the discount is 20% rather than 50%, applying to input, output, cached and reasoning tokens alike.

  • Meta · Muse Spark 1.2 standard official claim medium confidence

    Meta's frontier model is closed-weight and, since 2026-08-05, carries a published per-token price — the first time this catalog can compare against Meta's best model rather than only its open derivatives. Meta said on 2026-08-10 that it intends to open these weights "in the coming weeks"; that had not happened by the 2026-08-13 cutoff. No batch tier or long-context surcharge is published. CORRECTED 2026-08-27: context was recorded as 1,000,000. Meta's own models page states 1,048,576 for every Muse Spark ID, in a table cell and again in prose. The old figure understated the window by 4.6% on the field that decides whether a request is too large to send, so it was conservative rather than dangerous — but it was also the reason two pages of this site disagreed about the same model.

  • Meta · Muse Spark 1.2 contributor · Meta trains on your requests official claim medium confidence

    An opt-in tier at roughly a twentieth of the standard rate, priced on the condition that Meta trains future models on the requests. Carried because it is the cheapest frontier-lab tier in the catalog by a wide margin and therefore the hardest comparator any self-hosting case has to beat — but the calculator flags it, since data control is one of the main reasons to own hardware and this tier trades exactly that away. CORRECTED 2026-08-27: context was recorded as 1,000,000. Meta's own models page states 1,048,576 for every Muse Spark ID, in a table cell and again in prose. The old figure understated the window by 4.6% on the field that decides whether a request is too large to send, so it was conservative rather than dangerous — but it was also the reason two pages of this site disagreed about the same model.

  • Alibaba · Qwen3.8-Max standard · snapshot 0902 · international official claim medium confidence

    The 2.4T flagship previewed at WAIC on 2026-07-19 now has a published per-token price, so it is no longer excluded for want of one. Rates are the Singapore/international region; Alibaba's Beijing region prices it lower ($1.65 / $4.951) and this catalog records the international figures throughout. Max input is 991,808 tokens with 131,072 output inside the stated 1M window. The consolidated Model Studio pricing table did not yet carry this row at the cutoff, so the per-model documentation page is the source; confidence is medium accordingly. CORRECTED 2026-08-21: no cache-write price was recorded, so the engine fell back to the plain input rate of $2.00. Alibaba publishes "Explicit Cache Creation | 2.5 | Per 1M tokens" for the Singapore region — 125% of input, 25% above the fallback. Worth knowing that Alibaba documents three cache prices rather than two: explicit creation at 125% of input, an implicit hit at 12.5% and an explicit read at 8.5%. The figure recorded here as cachedInput is the printed implicit-hit one, $0.25; the explicit read is $0.17. CORRECTED 2026-08-29: two of those three percentages read 10% and 20%, which are neither the stored figures nor Alibaba's. Re-read from the same page, the Singapore table prints Input 2, Output 6, Input(Implicit Cache) 0.25, Explicit Cache Creation 2.5 and Explicit Cache Read 0.17, so the ratios are 12.5% and 8.5%; the Beijing column gives 0.206, 2.063 and 0.137 against an input of 1.65, which is the same three ratios. Every stored dollar figure was right and only the prose was wrong, which is the harder version of this defect to notice. SNAPSHOT 0902, read 2026-09-08: a dated build exists as API id qwen3.8-max-0902 (alias qwen3.8-max-2026-09-02), post-trained for engineering-scale projects and multi-tool agent orchestration. Alibaba and OpenRouter both state its pricing is unchanged from the base alias — same $2/$6, same three-tier cache pricing, same 1,000,000-token window — so it is named here rather than carried as a second row of identical numbers. Its release date is unresolved: TechNode's article is dated 2026-09-02 and OpenRouter's model page says September 3; no Alibaba-dated announcement was found to break the tie, and the two are recorded rather than averaged. The international price is also not printed in dollars on Alibaba's own page: that table renders in CNY (14.988 / 44.965 / 1.874 / 18.736 / 1.274 per 1M), and every one of those five divides by the same ~7.495 into the USD figures here, which OpenRouter states directly. TOKEN LIMITS, recorded 2026-09-13: maxInputTokens 991,808 and maxOutputTokens 131,072, re-read from the Context Limits table on Alibaba's qwen3.8-max page, which prints Max Input Length 991808, Max Output Length 131072 and Context Window 1000000. The same table gives thinking mode a lower Max Input Length of 983616, which a scenario cannot select and is not stored, so a thinking-mode request with more than 983,616 input tokens, up to 991,808, passes here although Alibaba refuses it. Whether reasoning counts toward the output length is not stated and not recorded; the table's Max Chain-of-Thought Length of 262144 is larger than the output length, which suggests reasoning is counted apart from it.

  • DeepSeek · DeepSeek V4 Flash standard official claim high confidence

    Hit/miss pricing: cache hits $0.0028/M, misses at standard input; no separate write fee, so writes are encoded at standard input. This flat rate ended at 16:00 UTC on 2026-08-16, replaced by peak/off-peak billing that is more expensive at every hour — see the successor tier. RETIRED 2026-08-21: the end date is now behind the research cutoff, so the tier is marked superseded. It stays resolvable, because scenario URLs shared while it was live must keep opening on the price they were shared at. REPOINTED 2026-09-11: supersededBy named deepseek-v4-flash-api-tod, which has itself now been retired behind DeepSeek-V4.1-Flash. A chain of retired entries leaves the UI offer to switch to the replacement pointing at something equally retired, so this now names the live tier directly.

  • DeepSeek · DeepSeek V4 Pro standard official claim high confidence

    Hit/miss pricing: cache hits $0.003625/M, misses at standard input; no separate write fee, so writes are encoded at standard input. This flat rate ended at 16:00 UTC on 2026-08-16, replaced by peak/off-peak billing that is more expensive at every hour — see the successor tier. RETIRED 2026-08-21: the end date is now behind the research cutoff, so the tier is marked superseded. It stays resolvable, because scenario URLs shared while it was live must keep opening on the price they were shared at.

  • DeepSeek · DeepSeek V4 Flash peak / off-peak · from 2026-08-16 official claim high confidence

    Effective 16:00 UTC on 2026-08-16. DeepSeek: "Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak)." CORRECTED 2026-08-27: this note previously quoted that sentence without "Monday through Friday" and derived 7 peak hours against 17 off-peak, which reads as 29.2% of the week at the peak rate. Weekends are entirely off-peak, so the real split is 35 peak hours of 168 (20.8%) against 133 off-peak (79.2%). Peak exposure therefore falls by two sevenths. The cost effect is much smaller than that sounds and is worth stating separately, because off-peak is half of peak rather than free: the blended multiplier moves from 0.7083x0.5 + 0.2917 = 0.6458 to 0.7917x0.5 + 0.2083 = 0.6042, so a workload spread evenly over the week costs 6.5% less, not 29% less. An earlier version of this note said 29%, which confused a change in exposure with a change in price. The clipped clause was the load-bearing one: everything else in the sentence is a schedule, and that phrase is what turns it into a weekly fraction. Listed figures are the PEAK rates; off-peak is exactly half on all three categories. READ THIS BEFORE TREATING IT AS A DISCOUNT: off-peak is not a discount off today's price, it is a discount off a newly raised one. Against the flat rate it replaces, off-peak is 1.6x on cache-miss input, 2.4x on output and 2.5x on cache hits; peak is 3.1x, 4.7x and 5.0x. The whole rate card went up several-fold and the peak/off-peak split is layered on top. RETIRED 2026-09-11: the pricing page now states that the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted but the corresponding models have been retired, their requests served by DeepSeek-V4.1-Flash and billed at the Flash price. The endpoint still answers, so this is not a 404 — it is the same name in front of a different model at a different price, which is the case a superseded tier exists for. It stays resolvable, because scenario URLs shared while it was live must keep opening on the price they were shared at.

  • DeepSeek · DeepSeek V4 Pro peak / off-peak · from 2026-08-16 official claim high confidence

    Effective 16:00 UTC on 2026-08-16, alongside the V4 Pro GA release. Listed figures are the PEAK rates; off-peak is exactly half on all three categories. The increase is steepest here: against the flat rate it replaces, off-peak is 1.5x on cache-miss input, 2.3x on output and 6.1x on cache hits, and peak reaches 3.0x, 4.6x and 12.1x. A cache-heavy agentic workload is hit hardest, since the cache-hit rate rose furthest. The peak/off-peak schedule is the same one quoted on the V4-Flash time-of-day tier, weekdays only; that entry carries the vendor's sentence and the weekly split. CONFLICT, recorded 2026-09-13 and deliberately left unresolved: DeepSeek's own pages disagree about this endpoint from 2026-09-14. The V4.1-Flash release note, dated 2026-09-10, says "Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches." The pricing page, read on 2026-09-13, footnotes the deepseek-v4-pro column with "In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." The footnote carries no date of its own, and neither page mentions the other. The tier stays current at the V4 Pro rates, which is what the endpoint bills until 2026-09-14 on both readings and after it on the pricing page's. No effectiveUntil is set, because that field would record the release note's date as the end of this rate and settle the conflict in the release note's favour. If the routing does take effect, a deepseek-v4-pro request is billed at the rates on deepseek-v4-1-flash-api-tod, which is a different model. Revisit after 2026-09-14. OUTPUT CAP, recorded 2026-09-13: 393,216 tokens. The pricing page prints MAX OUTPUT as “MAXIMUM: 384K” in a row shared by both model columns, and DeepSeek's chat completion reference spells it out: max_tokens “must be between 1 and 384K (393216)”. That reference does not say whether thinking-mode reasoning counts toward max_tokens, so that is not recorded, and the pricing page publishes no input limit below the window.

  • Alibaba · Qwen3.7-Max standard official claim medium confidence

    CONFLICT RESOLVED at the 2026-08-01 refresh. The earlier reading — that "$2.5, limited-time 50% off" implied a $5/$15 list price — was wrong: the list price on Alibaba's page is itself $2.50 in / $7.50 out, and the 50% discount applies on top of it, giving an effective $1.25/$3.75 that matches the third-party trackers. The discount badge appears only on the mutable `qwen3.7-max` alias; pinned dated snapshots such as `qwen3.7-max-2026-06-08` bill the plain $2.50/$7.50. This entry records the undiscounted list rate, which is what a pinned model id costs and what remains once the promotion ends. Cache hits bill at 10% of input and cache writes at 125%, both encoded. Context window is not stated on the pricing page, and was omitted rather than guessed. ADDED 2026-09-13: Alibaba's Model Studio page for qwen3.7-max prints a Context Limits table, Max Input Length 991808, Max Output Length 131072 and Context Window 1000000, and all three are now recorded. The same table gives thinking mode a lower Max Input Length of 983616, which a scenario cannot select and is not stored, so a thinking-mode request with more than 983,616 input tokens, up to 991,808, passes here although Alibaba refuses it. Whether reasoning counts toward the output length is not stated and not recorded; the table's Max Chain-of-Thought Length of 262144 is larger than the output length, which suggests reasoning is counted apart from it.

  • Alibaba · Qwen3.7-Plus standard official claim medium confidence

    Tiered pricing verified 2026-07-24: input $0.40 to 256K and $1.20 above, output $1.60 to 256K and $4.80 above (both 3x). No cached-input price encoded. Context window is not stated on the pricing page, and was omitted rather than guessed. ADDED 2026-09-13: Alibaba's Model Studio page for qwen3.7-plus prints a Context Limits table, Max Input Length 991808, Max Output Length 131072 and Context Window 1000000, and all three are now recorded. The same table gives thinking mode a lower Max Input Length of 983616, which a scenario cannot select and is not stored, so a thinking-mode request with more than 983,616 input tokens, up to 991,808, passes here although Alibaba refuses it. Whether reasoning counts toward the output length is not stated and not recorded; the table's Max Chain-of-Thought Length of 262144 is larger than the output length, which suggests reasoning is counted apart from it.

  • Xiaomi · MiMo V2.5 standard official claim medium confidence

    Overseas USD pricing; digits identical to DeepSeek V4 Flash on both tiers — apparent deliberate price-matching, verified on both primary pages 2026-07-24. Context from the model card; the pricing page does not state windows. Hit/miss model; writes at standard input. CORRECTED 2026-08-21: recorded as a hit/miss tier with writes priced at the standard input rate. Xiaomi does publish a distinct cache-write line, and it currently reads "Cache Write: Limited-time Free" — so the write bucket costs nothing here. Carried as zero with the caveat that it is a promotion with no published end date, which is a different kind of zero from OpenAI’s pre-5.6 structural one: this can be withdrawn without the model changing. OUTPUT CAP, recorded 2026-09-13: 131,072 tokens, reasoning included. Xiaomi's MiMo-V2.5 model page lists a max output of 128K tokens, and its chat API reference gives max_completion_tokens a required range of [1, 131072] and describes it as “An upper bound for the number of tokens that can be generated for a completion, including visible output tokens and reasoning tokens.”

  • Xiaomi · MiMo V2.5 Pro standard official claim medium confidence

    Cached input printed as $0.0036/M on MiMo's pricing page versus DeepSeek V4 Pro's $0.003625/M for the equivalent tier — a small cross-vendor print difference, preserved rather than reconciled. Context from the model card; the pricing page does not state windows. Hit/miss model; writes at standard input. CORRECTED 2026-08-21: recorded as a hit/miss tier with writes priced at the standard input rate. Xiaomi does publish a distinct cache-write line, and it currently reads "Cache Write: Limited-time Free" — so the write bucket costs nothing here. Carried as zero with the caveat that it is a promotion with no published end date, which is a different kind of zero from OpenAI’s pre-5.6 structural one: this can be withdrawn without the model changing. OUTPUT CAP, recorded 2026-09-13: 131,072 tokens, reasoning included. Xiaomi's MiMo-V2.5-Pro model page lists a max output of 128K tokens, and its chat API reference gives max_completion_tokens a required range of [1, 131072] and describes it as “An upper bound for the number of tokens that can be generated for a completion, including visible output tokens and reasoning tokens.”

  • Mistral AI · Mistral Medium 3.5 standard official claim high confidence

    Current proprietary flagship (v26.04); also published as open weights (Modified MIT). Generic 90% cached-input discount; no cache-write price published. TOKEN LIMITS, read 2026-09-13: no input limit or output cap separate from the window is recorded. Mistral's model page gives Mistral Medium 3.5's context as 256k, and its chat completion reference bounds a request only by that window: “The token count of your prompt plus max_tokens cannot exceed the model's context length.” The 256,000 stored is the decimal reading of 256k, because no page cited here spells this window out in digits. Mistral Large 3's card says 256k as well, and this catalog stores 262,144 for that model because the vLLM launch command on that card passes --max-model-len 262144; the card cited here carries no launch command. Under the binary reading, 262,144, a request of 256,001 to 262,144 tokens would be accepted by Mistral and is refused here.

  • Mistral AI · Mistral Medium 3.5 Batch · 50% multiplier estimated medium confidence

    Mistral publishes a generic 50% Batch discount and, separately, that cached input tokens reduce input cost by up to 90%. It never says whether the two stack, and publishes no per-model batch prices to derive it from. DOWNGRADED 2026-08-21 from high to medium confidence. Applying the batch multiplier to the cached bucket as well as the plain one is this catalog's reading, not a published rate — and it is a reading the other providers do not settle either way: Anthropic and OpenAI do halve every category, Google halves context caching on Gemini 3.7 Flash but not on 2.5 Pro, and xAI discounts 20% rather than 50%. A field graded high confidence on an unstated figure is the failure mode this catalog exists to avoid. REGRADED 2026-08-29: the 2026-08-21 pass moved confidence to medium and left evidenceKind at official-claim. Confidence and evidence kind are different axes — a reading is not the vendor's claim at any confidence — so this is now estimated. TOKEN LIMITS, read 2026-09-13: no input limit or output cap separate from the window is recorded. Mistral's model page gives Mistral Medium 3.5's context as 256k, and its chat completion reference bounds a request only by that window: “The token count of your prompt plus max_tokens cannot exceed the model's context length.” The 256,000 stored is the decimal reading of 256k, because no page cited here spells this window out in digits. Mistral Large 3's card says 256k as well, and this catalog stores 262,144 for that model because the vLLM launch command on that card passes --max-model-len 262144; the card cited here carries no launch command. Under the binary reading, 262,144, a request of 256,001 to 262,144 tokens would be accepted by Mistral and is refused here.

  • Z.ai · GLM-5.3 standard official claim high confidence

    Z.ai publishes this now, and the assumption it replaces was exact. Their pricing table reads "GLM-5.3 | $1.4 | $0.26 | Limited-time Free | $4.4" under the headings Input, Cached Input, Cached Input Storage and Output — identical to the GLM-5.2 row directly beneath it, which is what this tier had assumed on the vendor's same-base-model statement. UPGRADED 2026-08-21 from an assumption at low confidence. The previous note read "NOT A PUBLISHED PRICE" and described a model reachable only through a coding plan with no credits-per-token rate; that was true when written and is no longer. The id still says "assumed" because a scenario URL shared while it was one has to keep resolving — an id is a contract, not a description. Worth noting what Z.ai's third column is: a cache *storage* charge, currently waived, not a cache-write fee. There is no write line here, so the write bucket is priced at the plain input rate as it is for the other hit/miss providers — a convention of this catalog rather than a published fee. OUTPUT CAP, recorded 2026-09-13: 131,072 tokens. Z.ai's GLM-5.3 guide gives “a maximum output length of 128K tokens”, and its chat completion reference gives max_tokens a maximum of 131072 under the description “The maximum number of tokens for model output, the GLM-5.3, GLM-5.2, GLM-5.1, GLM-5, GLM-4.7, GLM-4.6 series supports 128K maximum output”. Neither page says whether reasoning counts toward max_tokens, so that is not recorded.

  • Z.ai · GLM-5.2 standard official claim high confidence

    Context window from Z.ai's GLM-5.2 guide, which gives it as 1M beside a maximum output of 128K. CORRECTED 2026-09-13: this sentence read “Context window confirmed at 1,048,576 tokens (128K max output) on Z.ai's own documentation at the 2026-08-01 refresh, having previously rested on secondary sources only” and cited no documentation page, and the guide, read on 2026-09-13, prints neither figure in digits. Z.ai spells out only the output figure: its chat completion reference gives the GLM-5.2 series “128K maximum output” with a maximum of 131072. The 1,048,576 applies the same binary reading to 1M, which Z.ai does not spell out. OUTPUT CAP, recorded 2026-09-13: maxOutputTokens holds that 131,072. Neither page says whether reasoning counts toward max_tokens, so that is not recorded. No long-context tiering — one flat rate across the window. Z.ai publishes no batch API, so there is no batch row to encode; cache storage is currently free as a limited-time promotion.

  • Z.ai · GLM-5.3-Flash standard · promotional through 2026-09-09 official claim high confidence

    Z.ai's pricing page carries this row with the list prices struck through: <del>$0.15</del> $0.075 input, <del>$0.03</del> $0.015 cached input, <del>$0.50</del> $0.25 output, and "GLM-5.3-Flash is available at a 50% discount (strikethrough prices are list prices). The promotion ends at 24:00 on September 9, 2026 (UTC+8, Singapore time)." A first extraction of the same page returned only the discounted figures with no sign a promotion existed, because the markup was dropped; the row was re-read from raw HTML to confirm. The standing rate is carried separately as glm-5-3-flash-api-list. No cache-write line is published, so the write bucket falls back to the plain input rate. Cached input storage is free for the same limited period. GLM-5.3 and GLM-5.2 carry no strikethrough and are unchanged. OUTPUT CAP, recorded 2026-09-13: 131,072 tokens. Z.ai's GLM-5.3-Flash guide lists Maximum Output Tokens as 128K, and its chat completion reference says “`GLM-5.3-Flash` supports a maximum output length of 128K” and gives max_tokens a maximum of 131072. Neither page says whether reasoning counts toward max_tokens, so that is not recorded. RETIRED 2026-09-14: the promotion ended at 24:00 on 2026-09-09 (UTC+8), which is behind the 2026-09-13 research cutoff, so the tier is marked superseded and points at glm-5-3-flash-api-list, the list rate that applies from 2026-09-10. The end date is the one Z.ai published and quoted above, so the retirement rests on no reading of the page after the cutoff. It stays resolvable, because scenario URLs shared while it was live must keep opening on the price they were shared at.

  • Z.ai · GLM-5.3-Flash list rate · from 2026-09-10 official claim high confidence

    The struck-through figures on Z.ai's own pricing row, which the same page names as the list prices the 50% promotion is measured against. Carried as its own tier so a horizon that outlives 2026-09-09 can be costed against the rate that will actually apply, rather than silently keeping a lapsed discount. OUTPUT CAP, recorded 2026-09-13: 131,072 tokens. Z.ai's GLM-5.3-Flash guide lists Maximum Output Tokens as 128K, and its chat completion reference says “`GLM-5.3-Flash` supports a maximum output length of 128K” and gives max_tokens a maximum of 131072. Neither page says whether reasoning counts toward max_tokens, so that is not recorded.

  • Alibaba · Qwen3.8-Flash standard · explicit cache official claim medium confidence

    Qwen Cloud's own model page: Input $0.15, Input (Implicit Cache) $0.016, Explicit Cache Creation $0.2, Explicit Cache Read $0.016, Output $0.47 per 1M tokens. The figure recorded as cachedInput is the printed implicit-hit rate, following the convention already documented on qwen3-8-max. Graded medium confidence because Qwen Cloud was the only primary source for these prices when they were recorded: Alibaba Model Studio's consolidated pricing table carried no qwen3.8-flash row, and its per-model page 302-redirected to a not-found page. That page resolves at the same address as of 2026-09-13 and is what the token limits below cite; its prices were not compared with these. This tier serves a different artefact from the open checkpoint — the model card calls it "the official version based on Qwen3.8-Flash-Next with more production features, e.g., 1M context length by default" — which is why the context here is 1M against the checkpoint's 262,144. Deliberately unscored: Artificial Analysis benchmarked the open weights, and its page for the served model returns 404, so carrying the checkpoint's 56 across would assert an equivalence the publisher's own wording does not support. TOKEN LIMITS, recorded 2026-09-13: maxInputTokens 991,808 and maxOutputTokens 131,072. Qwen Cloud's page rounds them to 991K and 131K; Alibaba Model Studio's page for qwen3.8-flash prints Max Input Length 991808, Max Output Length 131072 and Context Window 1000000. The same table gives thinking mode a lower Max Input Length of 983616, which a scenario cannot select and is not stored, so a thinking-mode request with more than 983,616 input tokens, up to 991,808, passes here although Alibaba refuses it. Whether reasoning counts toward the output length is not stated and not recorded; the table's Max Chain-of-Thought Length of 262144 is larger than the output length, which suggests reasoning is counted apart from it.

  • Anthropic · Claude Mythos 5 limited availability · 5m cache write official claim high confidence

    Priced identically to Claude Fable 5 on every category, which is why the two rows read the same: $10 base, $12.50 five-minute cache write, $20 one-hour cache write, $1 cache hit, $50 output, and $5/$25 on the Batch API, which claude-mythos-5-batch now carries. The one-hour tier is not encoded, as elsewhere in this catalog. Marked preview rather than current because the pricing table labels it “limited availability” and links it to anthropic.com/glasswing rather than to a model page. CORRECTED 2026-08-29: this note said Anthropic documents the 1M window at flat pricing “for it by name”, and excused the missing batch row as the practice “elsewhere in this catalog”. Neither held. The pricing page's long-context sentence names “Claude 4.6 and later models and Claude Mythos Preview”: Mythos 5 is covered by the version class, not by name, and Mythos Preview is a different model, named separately because it carries no version number. The document that does name it is the context-windows page, now cited. And batch was encoded in ten of this catalog's rows at the time, including for the identically-priced Fable 5. Uses the new-generation tokenizer (~30% more tokens for the same text than Sonnet 4.6 and earlier); raw per-token price comparisons across that boundary undercount effective cost. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included. Anthropic's thinking page lists each model's `max_tokens` ceiling and prints 128k for Claude Mythos 5. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Mythos 5 limited availability · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Mythos 5 Batch API · 5m cache write official claim high confidence

    Published, and not carried until now: the batch table prints “Claude Mythos 5 (limited availability) | $5 / MTok | $25 / MTok”, an exact 50% of the standard row, while the identically-priced Fable 5 already had a batch row here. A reader comparing deferrable work against this model was being shown twice what Anthropic charges for it. Preview for the same reason as the standard row. Uses the new-generation tokenizer (~30% more tokens for the same text than Sonnet 4.6 and earlier); raw per-token price comparisons across that boundary undercount effective cost. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included, and no higher cap on the Batch API. Anthropic's thinking page prints 128k for Claude Mythos 5 and a dash in its batches beta ceiling column, and its batch-processing page does not name Claude Mythos 5 among the models the output-300k-2026-03-24 header raises. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Mythos 5 Batch API · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Opus 5 fast mode · 5m cache write estimated high confidence

    Graded estimated rather than official-claim: two of the four figures here are read off Anthropic's table and two are computed by this catalog, and an entity carries one grade, so it takes the weaker. No capability score is carried either, for the reason qwen3-8-flash-api states about its own: fast mode is a different serving configuration, Artificial Analysis measured the standard one, and copying the index across would assert an equivalence nobody has published. Fast mode is a research preview that buys faster output at double the standard rate: $10 input and $50 output against Opus 5's $5 and $25, for the same model. The cache figures here are not published separately — Anthropic states that prompt-caching multipliers apply on top of fast-mode pricing, so the five-minute write is 1.25x and a hit is 0.1x of the fast base, giving $12.50 and $1. The premium applies across the full context window including requests over 200k tokens, so there is no long-context surcharge to encode. Three things this row cannot express: fast mode is unavailable on the Batch API, it is first-party Claude API only rather than AWS or partner clouds, and the data-residency multiplier stacks on top of it as well. Uses the new-generation tokenizer (~30% more tokens for the same text than Sonnet 4.6 and earlier); raw per-token price comparisons across that boundary undercount effective cost. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included, and read from the model rather than from the mode. Anthropic's thinking page prints 128k as Claude Opus 5's `max_tokens` ceiling. The fast-mode page publishes no ceiling of its own: fast mode is a `speed: "fast"` request to claude-opus-5 and “runs the same model with a faster inference configuration”, so the model's ceiling is the one recorded, and it “is not available with the Batch API”, so no batch ceiling applies. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Opus 5 fast mode · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • OpenAI · GPT-5.6 Cyber standard official claim high confidence

    The second of OpenAI's two Daybreak models, aliased as gpt-daybreak-red-latest; gpt-5.6-sol is the other, and this catalog already carries it. $12.50 input against Sol's $4.00 is 3.125x, and $75.00 against $20.00 is 3.75x. The cache-write figure is $15.625 as published, an exact 1.25x of base, and is carried unrounded because that is how the page prints it. CORRECTED 2026-08-29: this row carried no context window and argued that none could be known, because the pricing table prints four dashes where Sol has a long-context row. But the pricing table states no window for Sol either — OpenAI puts windows on the model pages, which is where Sol's own 1,050,000 came from, and gpt-5.6-cyber has one: 400,000, with 128,000 max output tokens. Reading the silence of one page as evidence of absence left the catalog's most expensive tier looking unbounded beside siblings recorded at 1,050,000. The two pages do disagree about one thing, and this row follows the pricing table. The model page carries the family's boilerplate line that prompts over 272K input tokens bill at 2x input and 1.5x output — the same sentence appears on gpt-5.4, gpt-5.5 and gpt-5.6-sol, all of which have 1,050,000-token windows and priced long-context columns. Cyber's own long-context columns are four dashes, so no long-context multiplier is encoded here; inventing $25.00 and $112.50 from a line that is not about this model would be a fabricated price. Batch is unavailable for Cyber: the batch table has no row for it. OpenAI warns that the gpt-daybreak-red-latest alias will be repointed at later frontier models with pricing adjusted to match, so this row describes gpt-5.6-cyber and not the alias. TOKEN LIMITS, recorded 2026-09-13: a maximum input of 272,000 tokens and an output cap of 128,000 inside the 400,000 window, from the model details on OpenAI's gpt-5.6-cyber page: “Maximum input tokens: 272,000” and “128,000 max output tokens”. That maximum input also means no prompt on this model can carry the “>272K input tokens” the boilerplate long-context line is about. Reasoning counts toward the output cap: OpenAI's token-counting guide says the `max_output_tokens` and `max_completion_tokens` parameters “limit all tokens generated by the model, including non-visible tokens.”

  • OpenAI · GPT-5.6 Cyber standard · data residency estimated low confidence

    Derived from an absence, and graded accordingly. The support table on OpenAI data residency page enumerates gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna and does not list gpt-5.6-cyber, while listing every one of its siblings. An argument from silence, so it is estimated rather than official-claim: the honest reading is that the endpoint does not serve this model, and the alternative is that the table is incomplete. Recorded as unavailable rather than left absent, because absent means nobody has checked and somebody has. Read 2026-09-10.

  • Google · Gemini 3.5 Flash-Lite standard official claim high confidence

    Google's cheapest current tier and, until now, absent from this catalog altogether, which left every comparison here priced against tiers a reader weighing high-volume work would not choose. $0.30 in, $2.50 out, $0.03 cached, one price for text, image, video and audio alike. Marked stable on Google's model list. No cache-write fee exists to encode: Google charges for cache reads and for storage, not for writes. Google also bills held prefixes at $1.00 per 1M tokens per hour, which this calculator has no concept of; a workload that keeps a large prefix warm across idle time is priced low here, and the cheaper the per-token rate the more that fee matters relative to it. No capability score: Artificial Analysis's page for this model was not read this session, and this catalog does not carry a score it has not seen dated. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-3.5-flash-lite. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Google · Gemini 3.5 Flash-Lite Batch / Flex official claim high confidence

    Prices stored rather than a multiplier, because the discount is not uniform and a single scalar would misprice a bucket. Input and output do halve, $0.30 to $0.15 and $2.50 to $1.25, but context caching goes $0.03 to $0.02, which is two thirds and not a half. Batch and Flex are published at the same figures. Cache storage stays at $1.00 per 1M tokens per hour and is not discounted at all. This is the same shape as gemini-2-5-pro-batch, where the cached rate does not move either. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-3.5-flash-lite. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Google · Gemini 3.1 Flash-Lite standard official claim high confidence

    The other Flash-Lite Google sells, and cheaper per input token than the 3.5 while charging more for it on audio: $0.25 for text, image and video against $0.50 for audio, and $0.025 against $0.05 cached. This calculator prices one token stream, so the text figure is what is stored and the audio premium is not modelled. $1.50 output. Marked stable on Google's model list. Google also bills held prefixes at $1.00 per 1M tokens per hour, which this calculator has no concept of; a workload that keeps a large prefix warm across idle time is priced low here, and the cheaper the per-token rate the more that fee matters relative to it. No capability score: Artificial Analysis's page for this model was not read this session, and this catalog does not carry a score it has not seen dated. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-3.1-flash-lite. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Google · Gemini 3.1 Flash-Lite Batch / Flex — 50% multiplier official claim high confidence

    Here a single multiplier is exactly right, unlike its 3.5 sibling: $0.125 input, $0.75 output and $0.0125 cached are each precisely half the standard figures, and Batch and Flex are published at the same rates. Cache storage halves too, $1.00 to $0.50 per 1M tokens per hour, which is the one Google discount of this kind the catalog has seen and still cannot express. TOKEN LIMITS, recorded 2026-09-13: an input token limit of 1,048,576 and an output token limit of 65,536, from Google's API model reference for gemini-3.1-flash-lite. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.”

  • Anthropic · Claude Fable 5.1 standard · 5m cache write official claim high confidence

    Same input, output and cache-write rates as Claude Fable 5, and one change: a cache read costs $0.25/M against Fable 5's $1.00/M. Anthropic states it as a 0.025x multiplier on base input where every other Claude model uses 0.1x, and frames the effect as roughly 25% cheaper for typical workloads and up to 45% for highly agentic ones — entirely from that one field. Fable 5 is not superseded: Anthropic's own pricing table lists both with no deprecation tag, and annotates genuinely retired models explicitly on the same table. The 2026-09-01 date is convergent press reporting; Anthropic's own announcement says only "September 2026". Uses the new-generation tokenizer (~30% more tokens for the same text than Sonnet 4.6 and earlier); raw per-token price comparisons across that boundary undercount effective cost. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included. Anthropic's thinking page lists each model's `max_tokens` ceiling and prints 128k for Claude Fable 5.1. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Fable 5.1 standard · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • Anthropic · Claude Fable 5.1 Batch API · 5m cache write official claim high confidence

    Official Batch API 50% multiplier. Same input, output and cache-write rates as Claude Fable 5, and one change: a cache read costs $0.25/M against Fable 5's $1.00/M. Anthropic states it as a 0.025x multiplier on base input where every other Claude model uses 0.1x, and frames the effect as roughly 25% cheaper for typical workloads and up to 45% for highly agentic ones — entirely from that one field. Fable 5 is not superseded: Anthropic's own pricing table lists both with no deprecation tag, and annotates genuinely retired models explicitly on the same table. The 2026-09-01 date is convergent press reporting; Anthropic's own announcement says only "September 2026". Uses the new-generation tokenizer (~30% more tokens for the same text than Sonnet 4.6 and earlier); raw per-token price comparisons across that boundary undercount effective cost. OUTPUT CAP, recorded 2026-09-13: 128,000 tokens, reasoning included, and no higher cap on the Batch API. Anthropic's thinking page prints 128k for Claude Fable 5.1 and a dash in its batches beta ceiling column, and its batch-processing page does not name Claude Fable 5.1 among the models the output-300k-2026-03-24 header raises. Anthropic's k is 1,000 in these limits: the same table prints as 300k the batch ceiling that the batch-processing page gives as 300,000. Thinking tokens “count toward `max_tokens` alongside the response text”, so reasoning counts toward the cap.

  • Anthropic · Claude Fable 5.1 Batch API · 5m cache write · data residency official claim high confidence

    Anthropic prices US-only inference at 1.1x "across all token pricing categories (input tokens, output tokens, cache writes, and cache reads)". The multiplier stacks with the others rather than replacing them: the caching section states it applies on top of the Batch API discount and of prompt-caching multipliers, and the fast-mode section says the same. Requested per request as inference_geo: "us"; global is the default and is what every rate in this entry is. Read 2026-09-10.

  • OpenAI · GPT-6 Astra standard official claim high confidence

    Read from OpenAI's own embedded pricing row, ["gpt-6-astra",10,1,12.5,50] for standard and ["gpt-6-astra",5,0.5,6.25,25] for batch — exactly half on all four columns. Above 272K input tokens the published long-context row is $20/$2/$25/$75, which is 2x on input, cached input and cache write but only 1.5x on output; the two multipliers are recorded separately for that reason. The 2026-09-03 release date is press reporting: openai.com's own announcement returned HTTP 403 to an automated fetch on both attempts, and the docs pages carry no date. TOKEN LIMITS, recorded 2026-09-13: a maximum input of 922,000 tokens and an output cap of 128,000 inside the 1,050,000 window, from the model details on OpenAI's gpt-6-astra page: “Maximum input tokens: 922,000” and “128,000 max output tokens”. Reasoning counts toward the output cap: OpenAI's token-counting guide says the `max_output_tokens` and `max_completion_tokens` parameters “limit all tokens generated by the model, including non-visible tokens.”

  • OpenAI · GPT-6 Astra standard · data residency official claim high confidence

    OpenAI states: "Data residency endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency." The support table on the same page lists this model as eligible, and every OpenAI tier in this catalog is priced from 2026-07-30 or later, so both conditions hold. Two things OpenAI does not say, and this catalog does not invent: which token categories the uplift covers - Anthropic enumerates input, output, cache reads and cache writes, OpenAI writes only "10% uplift" against the endpoint - and how it composes with the Batch and Flex discounts. Read 2026-09-10.

  • OpenAI · GPT-6 Astra Batch API official claim high confidence

    Published as its own row at exactly half the standard rate on all four columns rather than as a stated multiplier. Read from OpenAI's own embedded pricing row, ["gpt-6-astra",10,1,12.5,50] for standard and ["gpt-6-astra",5,0.5,6.25,25] for batch — exactly half on all four columns. Above 272K input tokens the published long-context row is $20/$2/$25/$75, which is 2x on input, cached input and cache write but only 1.5x on output; the two multipliers are recorded separately for that reason. The 2026-09-03 release date is press reporting: openai.com's own announcement returned HTTP 403 to an automated fetch on both attempts, and the docs pages carry no date. TOKEN LIMITS, recorded 2026-09-13: a maximum input of 922,000 tokens and an output cap of 128,000 inside the 1,050,000 window, from the model details on OpenAI's gpt-6-astra page: “Maximum input tokens: 922,000” and “128,000 max output tokens”. Reasoning counts toward the output cap: OpenAI's token-counting guide says the `max_output_tokens` and `max_completion_tokens` parameters “limit all tokens generated by the model, including non-visible tokens.”

  • OpenAI · GPT-6 Astra Batch API · data residency estimated medium confidence

    OpenAI states: "Data residency endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency." The support table on the same page lists this model as eligible, and every OpenAI tier in this catalog is priced from 2026-07-30 or later, so both conditions hold. Two things OpenAI does not say, and this catalog does not invent: which token categories the uplift covers - Anthropic enumerates input, output, cache reads and cache writes, OpenAI writes only "10% uplift" against the endpoint - and how it composes with the Batch and Flex discounts. Read 2026-09-10. Applied multiplicatively on top of this arrangement discount, which is a derivation: OpenAI states the uplift and the discount separately and never their interaction. Anthropic documents multiplicative composition for the equivalent case, which is why it is the reading taken, and is not a statement about OpenAI.

  • Google · Gemini 3.8 Flash standard · promotional through 2026-12-31 official claim high confidence

    Priced identically to Gemini 3.7 Flash on every published figure, promotional through 2026-12-31 and doubling on 2027-01-01. Google bills no cache write and instead bills cache storage by the hour, which nothing here models. CONTEXT WINDOW: 1,048,576, the input token limit on Google's API model reference for gemini-3.8-flash, which gives an output token limit of 65,536 beside it. CHANGED 2026-09-13 from 1,000,000, which this note called “a gap in the evidence rather than a smaller window” because “every Google page for 3.8 Flash states only "1M"”. Read on 2026-09-13, the model reference prints the digits; whether it did when that was written is not known. CORRECTED 2026-09-13: the note also pointed to “the DeepMind model card that gives 3.7's precise figure”. That card gives no precise figure. It says “a token context window of up to 1M”, as gemini-3-7-flash-card has recorded since 2026-08-29. The DeepMind Flash page for 3.8 lists Input tokens 1M and Output tokens 64k, the same two limits rounded. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. TOKEN LIMITS, recorded 2026-09-13: maxInputTokens and maxOutputTokens hold the two limits above. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.” Gemini 3.7 Flash is NOT superseded by this: Google's own announcement says 3.7 "remains fully supported for efficiency-first workloads", its model list tags both as stable side by side, and its deprecation schedule gives 3.7 no shutdown date.

  • Google · Gemini 3.8 Flash Batch / Flex · 50% multiplier · promotional through 2026-12-31 official claim high confidence

    Batch and Flex are published at exactly half the standard rate on input, output and context caching alike, so one multiplier is the right shape. Priced identically to Gemini 3.7 Flash on every published figure, promotional through 2026-12-31 and doubling on 2027-01-01. Google bills no cache write and instead bills cache storage by the hour, which nothing here models. CONTEXT WINDOW: 1,048,576, the input token limit on Google's API model reference for gemini-3.8-flash, which gives an output token limit of 65,536 beside it. CHANGED 2026-09-13 from 1,000,000, which this note called “a gap in the evidence rather than a smaller window” because “every Google page for 3.8 Flash states only "1M"”. Read on 2026-09-13, the model reference prints the digits; whether it did when that was written is not known. CORRECTED 2026-09-13: the note also pointed to “the DeepMind model card that gives 3.7's precise figure”. That card gives no precise figure. It says “a token context window of up to 1M”, as gemini-3-7-flash-card has recorded since 2026-08-29. The DeepMind Flash page for 3.8 lists Input tokens 1M and Output tokens 64k, the same two limits rounded. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. TOKEN LIMITS, recorded 2026-09-13: maxInputTokens and maxOutputTokens hold the two limits above. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.” Gemini 3.7 Flash is NOT superseded by this: Google's own announcement says 3.7 "remains fully supported for efficiency-first workloads", its model list tags both as stable side by side, and its deprecation schedule gives 3.7 no shutdown date.

  • Google · Gemini 3.8 Flash list rate · from 2027-01-01 official claim high confidence

    The rate Google publishes as taking effect 2027-01-01, exactly double the promotional rate on every category, so a horizon that outlives the promotion can be priced against what will actually be charged. Priced identically to Gemini 3.7 Flash on every published figure, promotional through 2026-12-31 and doubling on 2027-01-01. Google bills no cache write and instead bills cache storage by the hour, which nothing here models. CONTEXT WINDOW: 1,048,576, the input token limit on Google's API model reference for gemini-3.8-flash, which gives an output token limit of 65,536 beside it. CHANGED 2026-09-13 from 1,000,000, which this note called “a gap in the evidence rather than a smaller window” because “every Google page for 3.8 Flash states only "1M"”. Read on 2026-09-13, the model reference prints the digits; whether it did when that was written is not known. CORRECTED 2026-09-13: the note also pointed to “the DeepMind model card that gives 3.7's precise figure”. That card gives no precise figure. It says “a token context window of up to 1M”, as gemini-3-7-flash-card has recorded since 2026-08-29. The DeepMind Flash page for 3.8 lists Input tokens 1M and Output tokens 64k, the same two limits rounded. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. TOKEN LIMITS, recorded 2026-09-13: maxInputTokens and maxOutputTokens hold the two limits above. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.” Gemini 3.7 Flash is NOT superseded by this: Google's own announcement says 3.7 "remains fully supported for efficiency-first workloads", its model list tags both as stable side by side, and its deprecation schedule gives 3.7 no shutdown date.

  • Google · Gemini 3.8 Flash Batch / Flex · 50% multiplier · list rate from 2027-01-01 official claim high confidence

    Half the list rate on every category, from the same published table. Priced identically to Gemini 3.7 Flash on every published figure, promotional through 2026-12-31 and doubling on 2027-01-01. Google bills no cache write and instead bills cache storage by the hour, which nothing here models. CONTEXT WINDOW: 1,048,576, the input token limit on Google's API model reference for gemini-3.8-flash, which gives an output token limit of 65,536 beside it. CHANGED 2026-09-13 from 1,000,000, which this note called “a gap in the evidence rather than a smaller window” because “every Google page for 3.8 Flash states only "1M"”. Read on 2026-09-13, the model reference prints the digits; whether it did when that was written is not known. CORRECTED 2026-09-13: the note also pointed to “the DeepMind model card that gives 3.7's precise figure”. That card gives no precise figure. It says “a token context window of up to 1M”, as gemini-3-7-flash-card has recorded since 2026-08-29. The DeepMind Flash page for 3.8 lists Input tokens 1M and Output tokens 64k, the same two limits rounded. Google prints no combined window beside those two: its token guide defines the context window as “the combined limit of input and output tokens” and says to find the size on the models page, which gives these two limits. contextTokens holds the input limit, the smallest window that limit allows, so a request needing more than 1,048,576 tokens in total is refused here although Google may accept it. TOKEN LIMITS, recorded 2026-09-13: maxInputTokens and maxOutputTokens hold the two limits above. Whether thinking counts toward the output limit is not recorded: the one statement found, in Google's thinking guide, is about the Interactions API's `max_output_tokens` parameter, which it says sets “the maximum number of tokens a response can generate, including thought tokens.” Gemini 3.7 Flash is NOT superseded by this: Google's own announcement says 3.7 "remains fully supported for efficiency-first workloads", its model list tags both as stable side by side, and its deprecation schedule gives 3.7 no shutdown date.

  • Meta · Muse Spark 1.3 standard official claim high confidence

    Meta's pricing page states outright that both versions share the same standard pricing, and the Contributor tier carries over unchanged. Muse Spark 1.2 is not marked superseded: Meta calls 1.3 "the latest version ... Recommended for new work" without retiring 1.2, and prices them identically. The 2026-09-02 release date is secondary press; neither Meta docs page carries one.

    Meta ↗Meta ↗ read 2026-09-08
  • Meta · Muse Spark 1.3 contributor · Meta trains on your requests official claim high confidence

    An opt-in tier at roughly a twentieth of the standard rate, priced on the condition that Meta trains future models on the requests. Meta's pricing page states outright that both versions share the same standard pricing, and the Contributor tier carries over unchanged. Muse Spark 1.2 is not marked superseded: Meta calls 1.3 "the latest version ... Recommended for new work" without retiring 1.2, and prices them identically. The 2026-09-02 release date is secondary press; neither Meta docs page carries one.

    Meta ↗Meta ↗ read 2026-09-08
  • DeepSeek · DeepSeek V4.1 Flash peak / off-peak · from 2026-09-10 official claim high confidence

    Read from the pricing page on 2026-09-11 under the model column deepseek-flash, version DeepSeek-V4.1-Flash. Listed figures are the PEAK rates, the convention every DeepSeek tier here uses: cache-miss input $0.30, cache hit $0.006, output $1.20, and off-peak rates are half of the peak rates, with peak hours 01:00-04:00 and 06:00-10:00 UTC Monday through Friday and all other hours off-peak. Effective 04:00 UTC on 2026-09-10 per the release note. NO SEPARATE CACHE-WRITE FEE: the Context Caching guide describes caching as enabled by default for all users and documents no write charge, so a miss is billed at the standard input rate and persisted as a side effect — encoded here as cacheWriteUsdPerM equal to input, the same as every other DeepSeek tier. NO LONG-CONTEXT TIER: the rate is flat to the full 1M window, so there is no threshold and no inclusive-or-exclusive question. THIS IS CHEAPER THAN THE FLASH TIER IT REPLACES, which is unusual enough to state: against deepseek-v4-flash-api-tod it is 0.68x on cache-miss input, 0.91x on output and 0.43x on cache hits, at every hour. CORRECTED 2026-09-11: this sentence read "The concurrency limit rises from 500 to 2,500", which compared two different tiers and reported the comparison as a change over time. The pricing page publishes exactly two concurrency limits, one per model column: 2,500 against deepseek-flash, which is this tier, and 500 against deepseek-v4-pro. It carries no column for the retired V4-Flash tier at all, and neither the 2026-08-13 nor the 2026-09-10 release note states a limit for it, so there is no published figure for this tier to have risen from. What is publishable is the pair as it stands: 2,500 here, 500 on deepseek-v4-pro. The 2026-09-10 release note routes deepseek-v4-pro requests to this model at this price from 04:00 UTC on 2026-09-14 and says nothing about which limit they are then served under; the pricing page, read on 2026-09-13, says V4 Pro service continues after that date with its billing unchanged. deepseek-v4-pro-api-tod records the conflict. OUTPUT CAP, recorded 2026-09-13: 393,216 tokens. The pricing page prints MAX OUTPUT as “MAXIMUM: 384K” in a row shared by both model columns, and DeepSeek's chat completion reference spells it out: max_tokens “must be between 1 and 384K (393216)”. That reference does not say whether thinking-mode reasoning counts toward max_tokens, so that is not recorded, and the pricing page publishes no input limit below the window.

  • k3-dgx-b300-low SGLang v0.5.18 @ 71de97b2, TP8, native MXFP4 checkpoint estimated medium confidence

    The SGLang cookbook's B300 Low-Latency cell, marked Verified: TP8 on the native MXFP4 checkpoint with a bf16 cache and no speculation, 785 tok/s per GPU at concurrency 16 on requests of 8,192 prompt and 1,024 generated tokens. That figure counts both kinds of token, so the eight GPUs served 6,280 tok/s, 698 decode tok/s generated and 5,582 prefill tok/s of prompt. The cell's own medians close the loop: 16 streams at P50 TTFT 3,539 ms and P50 TPOT 19.47 ms generate 16 × 1,024 / (3.539 + 1,024 × 0.01947) = 697.9 tok/s. It is the low case because this recipe trades throughput for per-user speed, not because it reads the same run pessimistically. Decode and prefill split that one run's tokens at its own shape, and the calculator costs a scenario at whichever of the two limits it reaches first, which reproduces the run only at 8,192 tokens in and 1,024 out. At another prompt share and the same request length it understates, because one server can move GPU time between the two phases and the calculator holds each limit fixed. For requests longer than 9,216 tokens it overstates, because each prompt token then attends over, and each decode step reads, more cache than in the run. The run's random prompts shared no prefixes, and the calculator models no prefix cache: it charges every input token as prefill, so a workload that reuses prefixes gets more from this system than the calculator credits. CORRECTED 2026-09-13: this profile was 12,000 decode and 120,000 prefill tok/s, SGLang's 2,633 to 2,808 tok/s per GPU read as decode throughput on 8 GB300 and derated by 0.57, with prefill set at ten times decode. Those figures count prompt tokens as well and divide by the 16 or 32 GPUs of prefill/decode-disaggregated arms.

  • k3-dgx-b300-base SGLang v0.5.18 @ 71de97b2, TP8 + DCP8, native MXFP4 checkpoint estimated medium confidence

    The SGLang cookbook's B300 Balanced cell, marked Verified: TP8 with decode context parallelism across the eight GPUs, on the native MXFP4 checkpoint with a bf16 cache and no speculation, 1,395 tok/s per GPU at concurrency 64 on requests of 8,192 prompt and 1,024 generated tokens. That figure counts both kinds of token, so the eight GPUs served 11,160 tok/s, 1,240 decode tok/s generated and 9,920 prefill tok/s of prompt. At that shape both limits give 1.211 requests/s, and the cell's own medians agree: 64 streams at P50 TTFT 11,635 ms and P50 TPOT 40.19 ms complete 64 / (11.635 + 1,024 × 0.04019) = 1.212 requests/s, 1,241.5 generated tok/s. The cookbook publishes no Balanced point past concurrency 64 because the KDA state pool clamps admission at 101 running requests. Decode and prefill split that one run's tokens at its own shape, and the calculator costs a scenario at whichever of the two limits it reaches first, which reproduces the run only at 8,192 tokens in and 1,024 out. At another prompt share and the same request length it understates, because one server can move GPU time between the two phases and the calculator holds each limit fixed. For requests longer than 9,216 tokens it overstates, because each prompt token then attends over, and each decode step reads, more cache than in the run. The run's random prompts shared no prefixes, and the calculator models no prefix cache: it charges every input token as prefill, so a workload that reuses prefixes gets more from this system than the calculator credits. CORRECTED 2026-09-13: this profile was 21,000 decode and 210,000 prefill tok/s, SGLang's 2,633 tok/s per GPU read as decode throughput and multiplied by 8 GB300, with prefill set at ten times decode. That figure counts prompt tokens as well and divides by the 32 GPUs of a two-prefill, two-decode arm. The note also called the 3,000 tok/s planning guess carried before the weights existed low by roughly a factor of seven; that guess was 2.4 times this rate.

  • k3-dgx-b300-high SGLang v0.5.18 @ 71de97b2, TP8 + DCP8, DSPARK speculative decoding estimated low confidence

    The SGLang cookbook's B300 Balanced DSPARK cell: the Balanced recipe with DSPARK speculative decoding and --max-running-requests 256, 1,987 tok/s per GPU at concurrency 64 on requests of 8,192 prompt and 1,024 generated tokens. That figure counts both kinds of token, so the eight GPUs served 15,896 tok/s, 1,766 decode tok/s generated and 14,130 prefill tok/s of prompt, and the cell's medians agree: 64 × 1,024 / (12.038 + 1,024 × 0.02447) = 1,766.7 generated tok/s. Graded low because the cookbook pins the draft acceptance length with SGLANG_SIMULATE_ACC_LEN=4.5, so the cell reports what DSPARK's block of 7 delivers at that acceptance rather than an acceptance measured on this workload, and the cookbook asks readers to measure against the same recipe without speculation before adopting it. With DSPARK the KDA state pool clamps admission at 68 running requests. Speculation emits more than one token per decode step, so the bandwidth check, which assumes one, overstates what this rate demands of memory. Decode and prefill split that one run's tokens at its own shape, and the calculator costs a scenario at whichever of the two limits it reaches first, which reproduces the run only at 8,192 tokens in and 1,024 out. At another prompt share and the same request length it understates, because one server can move GPU time between the two phases and the calculator holds each limit fixed. For requests longer than 9,216 tokens it overstates, because each prompt token then attends over, and each decode step reads, more cache than in the run. The run's random prompts shared no prefixes, and the calculator models no prefix cache: it charges every input token as prefill, so a workload that reuses prefixes gets more from this system than the calculator credits. CORRECTED 2026-09-13: this profile was 22,500 decode and 225,000 prefill tok/s, SGLang's 2,808 tok/s per GPU read as decode throughput and scaled to 8 accelerators, with prefill set at ten times decode. That figure counts prompt tokens as well, divides by the 16 GPUs of a one-prefill, one-decode arm, and comes from a sweep tagged fp4 that the post never defines.

  • k3-mi355x-low SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV cache, measured on MI350X estimated low confidence

    The SGLang cookbook's MI350X Balanced cell at concurrency 16: TP8 on ROCm with AITER's A8W4 MoE kernels, Triton attention, an fp8_e4m3 cache where the calculator defaults to bf16, and no speculation, 462 tok/s per GPU on requests of 8,192 prompt and 1,024 generated tokens. That figure counts both kinds of token, so the eight GPUs served 3,696 tok/s, 411 decode tok/s generated and 3,285 prefill tok/s of prompt, and the cell's medians agree: 16 × 1,024 / (6.331 + 1,024 × 0.0328) = 410.4 generated tok/s. Measured on MI350X and carried to the MI355X, the part this catalog carries. The MI355X has the MI350X's 288 GB of HBM3E per GPU at 8 TB/s, by AMD's pages for both, and adds compute: 10.1 PFLOPs of peak MXFP4 performance against 9.2 and a 2,400 MHz peak engine clock against 2,200, at 1,400 W typical board power against 1,000, and the cookbook serves both from one recipe and one image. Decode is usually bounded by memory bandwidth, which the two share, and prefill by compute, where the MI355X has about a tenth more, so these figures more likely understate an MI355X than overstate it, by up to about a tenth where compute binds. Graded low because the cookbook marks every MI350X cell Final Verification In Progress and publishes no speed round for the MI355X. Decode and prefill split that one run's tokens at its own shape, and the calculator costs a scenario at whichever of the two limits it reaches first, which reproduces the run only at 8,192 tokens in and 1,024 out. At another prompt share and the same request length it understates, because one server can move GPU time between the two phases and the calculator holds each limit fixed. For requests longer than 9,216 tokens it overstates, because each prompt token then attends over, and each decode step reads, more cache than in the run. The run's random prompts shared no prefixes, and the calculator models no prefix cache: it charges every input token as prefill, so a workload that reuses prefixes gets more from this system than the calculator credits. UPGRADED 2026-09-13 from an assumption. The profile was 1,000 decode and 15,000 prefill tok/s, and its note read "Still a pure planning assumption. AMD published a day-0 Kimi K3 bring-up on MI355X but deliberately reported correctness only, with no throughput, TTFT or TPOT figures, so there is nothing to anchor these numbers to."

  • k3-mi355x-base SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV cache, measured on MI350X estimated low confidence

    The SGLang cookbook's MI350X Balanced cell at concurrency 64, on the recipe of the concurrency-16 cell: 813 tok/s per GPU on requests of 8,192 prompt and 1,024 generated tokens. That figure counts both kinds of token, so the eight GPUs served 6,504 tok/s, 723 decode tok/s generated and 5,781 prefill tok/s of prompt, and the cell's medians agree: 64 × 1,024 / (18.665 + 1,024 × 0.07039) = 722.2 generated tok/s. Measured on MI350X and carried to the MI355X, the part this catalog carries. The MI355X has the MI350X's 288 GB of HBM3E per GPU at 8 TB/s, by AMD's pages for both, and adds compute: 10.1 PFLOPs of peak MXFP4 performance against 9.2 and a 2,400 MHz peak engine clock against 2,200, at 1,400 W typical board power against 1,000, and the cookbook serves both from one recipe and one image. Decode is usually bounded by memory bandwidth, which the two share, and prefill by compute, where the MI355X has about a tenth more, so these figures more likely understate an MI355X than overstate it, by up to about a tenth where compute binds. Graded low because the cookbook marks every MI350X cell Final Verification In Progress and publishes no speed round for the MI355X. Decode and prefill split that one run's tokens at its own shape, and the calculator costs a scenario at whichever of the two limits it reaches first, which reproduces the run only at 8,192 tokens in and 1,024 out. At another prompt share and the same request length it understates, because one server can move GPU time between the two phases and the calculator holds each limit fixed. For requests longer than 9,216 tokens it overstates, because each prompt token then attends over, and each decode step reads, more cache than in the run. The run's random prompts shared no prefixes, and the calculator models no prefix cache: it charges every input token as prefill, so a workload that reuses prefixes gets more from this system than the calculator credits. UPGRADED 2026-09-13 from an assumption. The profile was 2,000 decode and 30,000 prefill tok/s, and its note read "Planning assumption, not a benchmark."

  • k3-mi355x-high SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV cache, DSPARK, measured on MI350X estimated low confidence

    The SGLang cookbook's MI350X Balanced DSPARK cell at concurrency 64, the same recipe with DSPARK speculative decoding: 913 tok/s per GPU on requests of 8,192 prompt and 1,024 generated tokens. That figure counts both kinds of token, so the eight GPUs served 7,304 tok/s, 812 decode tok/s generated and 6,492 prefill tok/s of prompt. This is the one cell whose medians do not close: 64 × 1,024 / (23.720 + 1,024 × 0.03145) = 1,171.9 generated tok/s, against 811.6 from the per-GPU figure. The profile takes the per-GPU figure, a whole-run total and the lower reading, while the same recipe's concurrency-16 DSPARK cell, 864 tok/s per GPU, agrees with its medians within 2%. Measured on MI350X and carried to the MI355X, the part this catalog carries. The MI355X has the MI350X's 288 GB of HBM3E per GPU at 8 TB/s, by AMD's pages for both, and adds compute: 10.1 PFLOPs of peak MXFP4 performance against 9.2 and a 2,400 MHz peak engine clock against 2,200, at 1,400 W typical board power against 1,000, and the cookbook serves both from one recipe and one image. Decode is usually bounded by memory bandwidth, which the two share, and prefill by compute, where the MI355X has about a tenth more, so these figures more likely understate an MI355X than overstate it, by up to about a tenth where compute binds. Graded low because the cookbook marks every MI350X cell Final Verification In Progress and publishes no speed round for the MI355X. Decode and prefill split that one run's tokens at its own shape, and the calculator costs a scenario at whichever of the two limits it reaches first, which reproduces the run only at 8,192 tokens in and 1,024 out. At another prompt share and the same request length it understates, because one server can move GPU time between the two phases and the calculator holds each limit fixed. For requests longer than 9,216 tokens it overstates, because each prompt token then attends over, and each decode step reads, more cache than in the run. The run's random prompts shared no prefixes, and the calculator models no prefix cache: it charges every input token as prefill, so a workload that reuses prefixes gets more from this system than the calculator credits. UPGRADED 2026-09-13 from an assumption. The profile was 4,000 decode and 60,000 prefill tok/s, and its note read "Optimistic planning assumption, not a benchmark."

  • k3-gb300-low SGLang v0.5.18, nine unified TP8 + DCP8 servers estimated low confidence

    Nine unified eight-GPU servers across the rack's 72 GPUs, each at the SGLang cookbook's B300 Balanced cell, 1,395 tok/s per GPU counting prompt and generated tokens: 9 × 8 × 1,395 = 100,440 tok/s, 11,160 decode tok/s generated and 89,280 prefill tok/s of prompt at 8,192 tokens in and 1,024 out. A GB300 carries the B300 GPU, and this case uses none of what the rack adds: no prefill/decode disaggregation and no NVLink domain wider than eight GPUs. The cell ran on a B300 1×8 node at concurrency 64, not on GB300, so it records no concurrency for the rack. Decode and prefill split that one run's tokens at its own shape, and the calculator costs a scenario at whichever of the two limits it reaches first, which reproduces the run only at 8,192 tokens in and 1,024 out. At another prompt share and the same request length it understates, because each server can move GPU time between the two phases and the calculator holds each limit fixed. For requests longer than 9,216 tokens it overstates, because each prompt token then attends over, and each decode step reads, more cache than in the run. The run's random prompts shared no prefixes, and the calculator models no prefix cache: it charges every input token as prefill, so a workload that reuses prefixes gets more from these servers than the calculator credits. CORRECTED 2026-09-13: this profile was 95,000 decode and 950,000 prefill tok/s, SGLang's 2,633 tok/s per GPU read as decode throughput on 8 GB300, multiplied by 72 at an assumed 50% scaling efficiency, with prefill set at ten times decode. That figure counts prompt tokens as well and divides by the 32 GPUs of a two-prefill, two-decode arm.

  • k3-gb300-base SGLang + Miles, two 2P:2D DCP8 arms and one TP8 + DCP8 server estimated low confidence

    Two copies of the DCP composition from SGLang's Kimi K3 post, two PP8 prefill workers feeding two DCP8 decode nodes on 32 GPUs at 2,633 tok/s per GPU, fill 64 of the rack's 72 GPUs, and one unified server at the cookbook's B300 Balanced cell, 1,395 tok/s per GPU, fills the other 8: 2 × 32 × 2,633 + 8 × 1,395 = 179,672 tok/s counting prompt and generated tokens, 19,964 decode tok/s generated and 159,708 prefill tok/s of prompt. Both runs split their tokens 8:1, the post's at 8,000 in and 1,000 out and the cookbook's at 8,192 and 1,024. The post's per-GPU figures divide by every GPU of an arm, prefill workers included, which the post does not state and its source record infers from the post's own measurements. Replication is the assumption: the post's 2P:2D arm at 2,633 sits within 2% of the single PP8 → DCP8 arm's 2,682.7, a gap its caption calls indistinguishable, but no published Kimi K3 serving run spans the rack, and the largest arm on the frontier uses 48 GPUs. The post prints no concurrency for either arm. The arms' prefill and decode workers, 64 of the rack's 72 GPUs, are sized for that 8:1 shape, and the calculator costs a scenario at whichever of the two limits it reaches first, which is what fixed workers give: a deployment serving another prompt share would re-divide the rack between prefill and decode workers and reach more than the calculator credits. For requests longer than the runs' 9,000 to 9,216 tokens it overstates, because each prompt token then attends over, and each decode step reads, more cache than in the runs. The post does not describe its prompts, and the calculator models no prefix cache: it charges every input token as prefill, so a workload that reuses prefixes gets more from this rack than the calculator credits. CORRECTED 2026-09-13: this profile was 130,000 decode and 1,300,000 prefill tok/s, the same 2,633 tok/s per GPU read as decode throughput on 8 GB300, multiplied by 72 at an assumed 70% scaling efficiency, with prefill set at ten times decode.

  • k3-gb300-high SGLang + Miles, four 1P:1D PP8 → TP8 arms and one TP8 + DCP8 server estimated low confidence

    Four copies of the highest untagged arm on SGLang's Kimi K3 serving frontier, one PP8 prefill worker feeding one TP8 decode node on 16 GPUs at 2,715 tok/s per GPU, fill 64 of the rack's 72 GPUs, and one unified server at the cookbook's B300 Balanced cell, 1,395 tok/s per GPU, fills the other 8: 4 × 16 × 2,715 + 8 × 1,395 = 184,920 tok/s counting prompt and generated tokens, 20,547 decode tok/s generated and 164,373 prefill tok/s of prompt. The arm of the same topology tagged fp4 peaks higher, at 2,808, and would make this 190,872; it is left out because the post never defines fp4, and a rate does not carry across a precision the catalog cannot name. The post's per-GPU figures divide by every GPU of an arm, prefill workers included, which the post does not state and its source record infers from the post's own measurements. Replication is the assumption, and a larger one than in the base case: four copies where the post measured at most two, and no published Kimi K3 serving run spans the rack. The arms' prefill and decode workers, 64 of the rack's 72 GPUs, are sized for that 8:1 shape, and the calculator costs a scenario at whichever of the two limits it reaches first, which is what fixed workers give: a deployment serving another prompt share would re-divide the rack between prefill and decode workers and reach more than the calculator credits. For requests longer than the runs' 9,000 to 9,216 tokens it overstates, because each prompt token then attends over, and each decode step reads, more cache than in the runs. The post does not describe its prompts, and the calculator models no prefix cache: it charges every input token as prefill, so a workload that reuses prefixes gets more from this rack than the calculator credits. CORRECTED 2026-09-13: this profile was 180,000 decode and 1,800,000 prefill tok/s, SGLang's 2,808 tok/s per GPU read as decode throughput on 8 GB300, multiplied by 72 at an assumed 90% scaling efficiency, with prefill set at ten times decode. That figure counts prompt tokens as well and divides by the 16 GPUs of a one-prefill, one-decode arm.

  • 8×H200 · Llama 2 70B MLPerf Inference v5.0 (server) measured high confidence
  • 8×MI325X · Llama 2 70B MLPerf Inference v5.0 (server) measured high confidence
  • DGX B200 · Llama 2 70B MLPerf Inference v5.0 preview (server) measured medium confidence
  • 8×H200 · Mixtral 8x7B MLPerf Inference v4.1 (server) measured high confidence
  • 8×H100 · Mixtral 8x7B MLPerf Inference v4.1 (server) measured high confidence
  • MI300X server · Llama 2 70B MLPerf Inference v4.1 (ROCm) (server) measured high confidence
  • H100 server · Llama 2 70B MLPerf Inference v4.1 (TensorRT) (server) measured high confidence
  • MI300X server · Llama 2 70B MLPerf Inference v4.1 (ROCm) (offline) measured high confidence
  • H100 server · Llama 2 70B MLPerf Inference v4.1 (TensorRT) (offline) measured high confidence
  • 8×B300 Cisco UCS C880A M8 · DeepSeek-R1 MLPerf Inference v6.0 (offline) measured high confidence
  • 8×B300 Cisco UCS C880A M8 · DeepSeek-R1 MLPerf Inference v6.0 (server) measured high confidence
  • GB300 NVL72 (full 72-GPU rack) · DeepSeek-R1 MLPerf Inference v6.0 (offline) measured high confidence
  • GB300 NVL72 (full 72-GPU rack) · DeepSeek-R1 MLPerf Inference v6.0 (server) measured high confidence
  • GB300 NVL72 (full 72-GPU rack) · GPT-OSS-120B MLPerf Inference v6.0 (offline) measured high confidence
  • GB300 NVL72 (full 72-GPU rack) · GPT-OSS-120B MLPerf Inference v6.0 (server) measured high confidence
  • GB200 NVL72, 64 of 72 GPUs (CoreWeave) · DeepSeek-R1 MLPerf Inference v6.0 (offline) measured high confidence
  • 8×MI355X · Llama 2 70B MLPerf Inference v6.0 (ROCm, FP4) (offline) measured high confidence
  • 8×MI355X · Llama 2 70B MLPerf Inference v6.0 (ROCm, FP4) (interactive) measured high confidence
  • 8×MI300X Supermicro · Llama 2 70B MLPerf Inference v5.1 (offline) measured high confidence
  • 8×MI300X Supermicro · Llama 2 70B MLPerf Inference v5.1 (interactive) measured high confidence
  • 8×MI325X QuantaGrid · Mixtral 8x7B MLPerf Inference v5.1 (offline) measured high confidence
  • 8×B200 · Llama 3.1 405B MLPerf Inference v5.1 (offline) measured high confidence
  • B200 (GPU count not disclosed on the compare view) · Qwen 3.5 397B-A17B InferenceX dashboard (FP8 default) (69 tok/s/user interactivity point) estimated medium confidence
  • H200 · Qwen 3.5 397B-A17B InferenceX dashboard (FP8 default) (69 tok/s/user interactivity point) estimated medium confidence
  • 2×GB300 (single node) · DeepSeek-V3.2-NVFP4 vLLM v0.14.1, CUDA 13.0 (mixed-context (ISL 2K/OSL 1K), TP2) measured medium confidence
  • GB300 (TP8) · Kimi K3 (MXFP4) vLLM day-0 build, FP8 KV cache (single user (batch 1), 8K in / 1K out) vendor reported medium confidence

    Published by the vLLM team about its own engine on day zero. A real measurement, but self-reported by the party with an interest in the result, and not yet independently reproduced.

  • GB300 (TP16) · Kimi K3 (MXFP4) vLLM day-0 build, FP8 KV cache (single user (batch 1), 8K in / 1K out) vendor reported medium confidence

    Published by the vLLM team about its own engine on day zero; not independently reproduced.

  • 16× GB300: 1 PP8 prefill worker → 1 TP8 decode node (1P:1D) · Kimi K3 SGLang + Miles, prefill/decode disaggregated, fp4 arm (serving frontier at its throughput end, 8K in / 1K out, concurrency 1024) vendor reported medium confidence

    Published by the SGLang/LMSYS team about its own engine on day zero. Reproducing it requires running the same disaggregated prefill/decode split, not a default single-server deployment.

  • 16× GB300: 1 PP8 prefill worker → 1 TP8 decode node (1P:1D) · Kimi K3 SGLang + Miles, prefill/decode disaggregated (serving frontier, untagged PP8 → TP8 arm at its peak, 8K in / 1K out) vendor reported medium confidence

    Published by the SGLang/LMSYS team about its own engine on day zero. Reproducing it requires running the same disaggregated prefill/decode split, not a default single-server deployment. The figure is read from the chart's data coordinates rather than quoted from its text.

  • 32× GB300: 2 PP8 prefill workers → 2 DCP8 decode nodes (2P:2D) · Kimi K3 SGLang + Miles, prefill/decode disaggregated, TP8 + DCP8 decode (serving frontier, the DCP composition at its peak, 8K in / 1K out) vendor reported medium confidence

    Published by the SGLang/LMSYS team about its own engine on day zero. Reproducing it requires running the same disaggregated prefill/decode split, not a default single-server deployment.

  • 8× GB300 (2×4 trays), TP8 + DCP8 + host-memory KV tier · Kimi K3 SGLang + Miles, DCP8, hicache ratio 2 (AgentX replay of coding-agent sessions, 48 concurrent sessions) vendor reported low confidence

    Published by the SGLang/LMSYS team about its own engine; not independently reproduced. DOWNGRADED 2026-09-13 from medium: the figure the post cites for 541 tok/s implies about 427.

  • 8× B300 (1×8, prefill and decode unified) · Kimi K3 (MXFP4) SGLang v0.5.18 @ 71de97b2, TP8 (cookbook Low-Latency cell, random 8,192 in / 1,024 out, concurrency 16) vendor reported medium confidence

    Published by the SGLang project about its own engine; not independently reproduced.

  • 8× B300 (1×8, prefill and decode unified) · Kimi K3 (MXFP4) SGLang v0.5.18 @ 71de97b2, TP8 + DCP8 (cookbook Balanced cell, random 8,192 in / 1,024 out, concurrency 64) vendor reported medium confidence

    Published by the SGLang project about its own engine; not independently reproduced.

  • 8× B300 (1×8, prefill and decode unified) · Kimi K3 (MXFP4) SGLang v0.5.18 @ 71de97b2, TP8 + DCP8, DSPARK speculative decoding (cookbook Balanced DSPARK cell, random 8,192 in / 1,024 out, concurrency 64) vendor reported low confidence

    Published by the SGLang project about its own engine; not independently reproduced. Graded low because the acceptance length behind the figure is simulated.

  • 8× MI350X (1×8, prefill and decode unified) · Kimi K3 (MXFP4) SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV (cookbook Balanced cell, random 8,192 in / 1,024 out, concurrency 16) vendor reported low confidence

    Published by the SGLang project about its own engine, before the cookbook's final verification round on the released weights.

  • 8× MI350X (1×8, prefill and decode unified) · Kimi K3 (MXFP4) SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV (cookbook Balanced cell, random 8,192 in / 1,024 out, concurrency 64) vendor reported low confidence

    Published by the SGLang project about its own engine, before the cookbook's final verification round on the released weights.

  • 8× MI350X (1×8, prefill and decode unified) · Kimi K3 (MXFP4) SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV, DSPARK speculative decoding (cookbook Balanced DSPARK cell, random 8,192 in / 1,024 out, concurrency 64) vendor reported low confidence

    Published by the SGLang project about its own engine, before the cookbook's final verification round on the released weights.

  • Agentic coding Request shape and cache mix measured high confidence

    Token-weighted mean of 64,680 individual API requests across 912 threads recorded between 2026-06-10 and 2026-08-16. The input mix measured 97.97% cache hit, 1.99% cache write and 0.037% uncached; rounding to 98/2/0 is deliberate, because giving the uncached bucket a token 1% is LESS accurate than zero when uncached input costs ten times a cache hit. An independent Claude Code user publishing a 30-day aggregate lands at 96.58/3.35/0.07, within 1.4 points. Caveat worth carrying: this is one person’s working style, with unusually long sessions — public figures put typical cache-write shares at 3-7% rather than 2%, which is exactly what the session-length curve predicts.

  • Chat assistant Request shape and cache mix estimated medium confidence

    Derived, not measured. A conversation of N turns over a stable prefix bills 2/(N+1) of its input as cache writes; at a dozen turns that is 15%. Corroborated by the shortest real sessions in the recorded corpus, which measure 83.84% hit and 14.90% write. The token counts are judgement.

  • Long-document reading Request shape and cache mix assumption low confidence

    An assumption with its arithmetic stated: one document written to cache once and re-read for each of five follow-up questions gives a 20% write share. Change the question count and the mix moves a long way — two questions gives 50%, twenty gives 5% — so this is the number in the catalog most worth overriding. At a single question caching cannot pay for itself at all and the honest mix is none.

    Anthropic ↗ read 2026-08-16
  • Million-token context Request shape and cache mix assumption low confidence

    The same shape as long-document reading, assuming twelve questions rather than five because a million-token write is worth amortising further. Same caveat: the question count is the assumption doing the work.

    Anthropic ↗ read 2026-08-16
  • Output-heavy drafting Request shape and cache mix assumption low confidence

    A short brief against a shared style guide. Assumes roughly 60% of the prompt is a cacheable shared prefix. Output dominates the bill at this shape, so the input mix barely moves the answer — which is why a low-confidence assumption is tolerable here and would not be in the agentic preset.

    Anthropic ↗ read 2026-08-16
Inference Economics · static, local-first calculationCatalog 0.8.1 · cutoff 2026-09-13 · app v0.8.1 · build 1c759ca