Throughput

Every tokens-per-second figure this repository has found published, grouped by model and named down to the hardware, the scenario and the engine that produced it. These rows are evidence and not an input: the calculator costs a scenario from the modeled planning profiles instead, which the checkbox below reveals. Publishers count different tokens, so every figure says which it counts where that is known: on requests of 8,192 prompt tokens and 1,024 generated ones, a rate counting both is nine times a rate counting the generated tokens alone, and one user's decode speed is not a throughput at all. A figure is printed in the scope it was published in, the whole run or one accelerator. The other scope is derived only where the run's accelerator count is known, and the ≈ in front of it marks that arithmetic: nobody published that figure. Bars compare figures of one kind within a card, and every other row says its scale is not directly comparable.

Llama 2 70B

8×MI355X offline · MLPerf Inference v6.0 (ROCm, FP4)
aggregate 103,480 tok/s · ≈ 12,935 tok/s per accelerator (8× MI355X)
~12,935 tok/s per GPU; AMD states FP4 precision.
DGX B200 server · MLPerf Inference v5.0 preview
aggregate 92,338.7 tok/s · ≈ 11,542 tok/s per accelerator (8× B200)
Preview result; benchmark harness and latency constraints apply.
8×MI355X interactive · MLPerf Inference v6.0 (ROCm, FP4)
aggregate 73,608 tok/s · ≈ 9,201 tok/s per accelerator (8× MI355X)
~9,201 tok/s per GPU; v6.0's Interactive scenario applies a tighter per-token latency SLA.
8×MI325X server · MLPerf Inference v5.0
aggregate 30,724.5 tok/s · ≈ 3,841 tok/s per accelerator (8× MI325X)
Server result; benchmark harness and latency constraints apply.
8×MI300X Supermicro offline · MLPerf Inference v5.1
aggregate 27,804 tok/s · ≈ 3,476 tok/s per accelerator (8× MI300X)
Neutral-ledger complement to the AMD-blog row (22,021 tok/s) — different harness and latency methodology; kept separate, never averaged.
8×H200 server · MLPerf Inference v5.0
aggregate 27,417.3 tok/s · ≈ 3,427 tok/s per accelerator (8× H200)
Server result; benchmark harness and latency constraints apply.
MI300X server offline · MLPerf Inference v4.1 (ROCm)
aggregate 24,110 tok/s · ≈ 3,014 tok/s per accelerator (8× MI300X)
AMD's own MLPerf Inference v4.1 closed-division Offline result for 8×MI300X with EPYC Turin, 24,109.8 tokens per second in its submission log, against NVIDIA's DGX H100 at 24,524.9. CORRECTED 2026-09-13: this row was rocm-mi300x-llama31-405b, labelled Llama 3.1 405B and linked to that model, because AMD's page labels the table row that way. Round 4.1 had no Llama 3.1 405B benchmark, and both figures are the Llama 2 70B Offline results in the two submissions' own logs. On the throughput page the row sat beside an audited 405B result some fifteen times lower per accelerator.
MI300X server server · MLPerf Inference v4.1 (ROCm)
aggregate 22,021 tok/s · ≈ 2,753 tok/s per accelerator (8× MI300X)
MLPerf Inference v4.1 closed-division Server result, quoted on AMD's page: AMD's own 8×MI300X submission with EPYC Turin logs 22,020.92 completed tokens per second, against NVIDIA's DGX H100 at 21,605.79. CPU, runtime and latency caveats apply.
8×MI300X Supermicro interactive · MLPerf Inference v5.1
aggregate 8,840 tok/s · ≈ 1,105 tok/s per accelerator (8× MI300X)
~1,105 tok/s per GPU.
H100 server server · MLPerf Inference v4.1 (TensorRT)
reported 21,605 tok/s
NVIDIA's own MLPerf Inference v4.1 closed-division Server result for DGX H100, 21,605.79 completed tokens per second in its submission log, which AMD's page quotes as the comparison for its MI300X row. CORRECTED 2026-09-13: this note called the figure AMD's report, made without independent audit; it is NVIDIA's own submission, which AMD only quotes.
H100 server offline · MLPerf Inference v4.1 (TensorRT)
reported 24,525 tok/s
NVIDIA's own MLPerf Inference v4.1 closed-division Offline result for DGX H100, 24,524.9 tokens per second in its submission log, quoted on AMD's page as the comparison. CORRECTED 2026-09-13: this row was rocm-h100-llama31-405b, labelled Llama 3.1 405B and linked to that model, from the same mislabelled table row as its MI300X sibling, and its note called the figure AMD's report.

Mixtral 8x7B

8×MI325X QuantaGrid offline · MLPerf Inference v5.1
aggregate 68,781 tok/s · ≈ 8,598 tok/s per accelerator (8× MI325X)
A second 8×MI325X submission (ASUSTeK) reported 68,121 offline — close but distinct; both valid, not averaged.
8×H200 server · MLPerf Inference v4.1
aggregate 57,177 tok/s · ≈ 7,147 tok/s per accelerator (8× H200)
8×H100 server · MLPerf Inference v4.1
reported 50,796 tok/s

DeepSeek-R1

GB300 NVL72 (full 72-GPU rack) offline · MLPerf Inference v6.0
aggregate 673,936 tok/s · ≈ 9,360 tok/s per accelerator (72× GB300)
Nebius submission; ~9,360 tok/s per GPU. CoreWeave submitted the same hardware class separately without publishing exact figures.
8×B300 Cisco UCS C880A M8 offline · MLPerf Inference v6.0
aggregate 69,251.4 tok/s · ≈ 8,656 tok/s per accelerator (8× B300)
8-GPU node; ~8,656 tok/s per GPU. Submission precision not independently confirmed.
GB300 NVL72 (full 72-GPU rack) server · MLPerf Inference v6.0
aggregate 575,580 tok/s · ≈ 7,994 tok/s per accelerator (72× GB300)
~7,994 tok/s per GPU.
GB200 NVL72, 64 of 72 GPUs (CoreWeave) offline · MLPerf Inference v6.0
aggregate 468,665 tok/s · ≈ 7,323 tok/s per accelerator (64× GB200)
Scale mismatch flagged: 16 nodes/64 GPUs, not a fully populated 72-GPU rack; ~7,323 tok/s per GPU.
8×B300 Cisco UCS C880A M8 server · MLPerf Inference v6.0
aggregate 58,553.3 tok/s · ≈ 7,319 tok/s per accelerator (8× B300)
~7,319 tok/s per GPU; MLPerf latency-constrained server scenario.

GPT-OSS-120B

GB300 NVL72 (full 72-GPU rack) server · MLPerf Inference v6.0
aggregate 1,096,770 tok/s · ≈ 15,233 tok/s per accelerator (72× GB300)
~15,233 tok/s per GPU.
GB300 NVL72 (full 72-GPU rack) offline · MLPerf Inference v6.0
aggregate 1,046,150 tok/s · ≈ 14,530 tok/s per accelerator (72× GB300)
~14,530 tok/s per GPU; identical figures in NVIDIA's and Nebius's recaps.

Llama 3.1 405B

8×B200 offline · MLPerf Inference v5.1
aggregate 1,650 tok/s · ≈ 206 tok/s per accelerator (8× B200)
~206 tok/s per GPU; audited non-preview complement to the v5.0 preview Llama-2-70B row on the same hardware class.

Qwen 3.5 397B-A17B

B200 (GPU count not disclosed on the compare view) 69 tok/s/user interactivity point · InferenceX dashboard (FP8 default)
3,225.9 tok/s per accelerator
tok/s per GPU, interpolated between measured InferenceX runs (per the site's own methodology note); GPU count/TP not shown on the compare view. Which tokens it counts was not recorded: InferenceX derives three throughputs per GPU for every run, tput_per_gpu, input_tput_per_gpu and output_tput_per_gpu, and this row does not say which of them it was read from.
H200 69 tok/s/user interactivity point · InferenceX dashboard (FP8 default)
485.3 tok/s per accelerator
Same compare view as the B200 row; the ~6.6x gap is one interactivity slice, not a universal multiplier. Which of InferenceX's three per-GPU throughputs it is, total, input or output, was not recorded either.

DeepSeek-V3.2-NVFP4

2×GB300 (single node) mixed-context (ISL 2K/OSL 1K), TP2 · vLLM v0.14.1, CUDA 13.0
2,816 tok/s per accelerator
tok/s per GPU on a 2-GPU dev node — deliberately NOT linked to the 72-GPU rack entity. Prefill-only on the same node: 7,360 tok/s per GPU. The post calls the 2,816 figure output throughput, so it counts generated tokens.

Kimi K3 (MXFP4)model: Kimi K3

8× B300 (1×8, prefill and decode unified) cookbook Balanced DSPARK cell, random 8,192 in / 1,024 out, concurrency 64 · SGLang v0.5.18 @ 71de97b2, TP8 + DCP8, DSPARK speculative decoding
1,987 tok/s per accelerator · ≈ 15,896 tok/s across 8× B300
Marked Verified in the cookbook, with a simulated acceptance: P50 TTFT 12,038 ms, P50 TPOT 24.47 ms and 1,987 tok/s per GPU at concurrency 64, the Balanced recipe with DSPARK speculative decoding and --max-running-requests 256. The cookbook pins the draft acceptance length with SGLANG_SIMULATE_ACC_LEN=4.5, so the cell reports what DSPARK's block of 7 delivers at that acceptance, not an acceptance rate measured on this workload, and it asks readers to measure against the same recipe without speculation before adopting it. With DSPARK the KDA state pool caps admission at 68 running requests. The per-GPU figure counts prompt and generated tokens: 64 × 1,024 / (12.038 + 1,024 × 0.02447) = 1,766.7 generated tok/s, against 1,987 × 8 × 1,024 / 9,216 = 1,766.2.
8× B300 (1×8, prefill and decode unified) cookbook Balanced cell, random 8,192 in / 1,024 out, concurrency 64 · SGLang v0.5.18 @ 71de97b2, TP8 + DCP8
1,395 tok/s per accelerator · ≈ 11,160 tok/s across 8× B300
Marked Verified in the cookbook: P50 TTFT 11,635 ms, P50 TPOT 40.19 ms and 1,395 tok/s per GPU at concurrency 64, with the native MXFP4 checkpoint, decode context parallelism across the eight GPUs, a bf16 cache and no speculation. The per-GPU figure counts prompt and generated tokens: 64 × 1,024 / (11.635 + 1,024 × 0.04019) = 1,241.5 generated tok/s, against 1,395 × 8 × 1,024 / 9,216 = 1,240.0. The KDA state pool caps this recipe at 101 running requests, which is why the cookbook publishes no Balanced point past concurrency 64. Measured with --random-range-ratio 1.0, --warmup-requests 64 and --flush-cache.
8× MI350X (1×8, prefill and decode unified) cookbook Balanced DSPARK cell, random 8,192 in / 1,024 out, concurrency 64 · SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV, DSPARK speculative decoding
913 tok/s per accelerator
Marked Final Verification In Progress in the cookbook: P50 TTFT 23,720 ms, P50 TPOT 31.45 ms and 913 tok/s per GPU at concurrency 64, with DSPARK speculative decoding on the shared MI350X/MI355X recipe. Here the medians do not reproduce the per-GPU figure: 64 × 1,024 / (23.720 + 1,024 × 0.03145) = 1,171.9 generated tok/s, against 913 × 8 × 1,024 / 9,216 = 811.6, a ratio of 0.69, while the same recipe's concurrency-16 DSPARK cell, 864 tok/s per GPU, agrees within 2%. The per-GPU figure is a whole-run total and the medians describe a typical request; the cookbook says neither which to trust nor whether this cell's acceptance length is simulated as the B300 cells' is. Not linked to a catalog system, which carries the MI355X and not the MI350X.
8× MI350X (1×8, prefill and decode unified) cookbook Balanced cell, random 8,192 in / 1,024 out, concurrency 64 · SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV
813 tok/s per accelerator
Marked Final Verification In Progress in the cookbook: P50 TTFT 18,665 ms, P50 TPOT 70.39 ms and 813 tok/s per GPU at concurrency 64, on the same recipe as the concurrency-16 cell. The per-GPU figure counts prompt and generated tokens: 64 × 1,024 / (18.665 + 1,024 × 0.07039) = 722.2 generated tok/s, against 813 × 8 × 1,024 / 9,216 = 722.7. Not linked to a catalog system, which carries the MI355X and not the MI350X.
8× B300 (1×8, prefill and decode unified) cookbook Low-Latency cell, random 8,192 in / 1,024 out, concurrency 16 · SGLang v0.5.18 @ 71de97b2, TP8
785 tok/s per accelerator · ≈ 6,280 tok/s across 8× B300
Marked Verified in the cookbook: P50 TTFT 3,539 ms, P50 TPOT 19.47 ms and 785 tok/s per GPU at concurrency 16, with the native MXFP4 checkpoint, no --kv-cache-dtype flag, so the cache stays at the model's bf16, and no speculation. The per-GPU figure counts prompt and generated tokens: 16 streams at those medians generate 16 × 1,024 / (3.539 + 1,024 × 0.01947) = 697.9 tok/s, and 785 × 8 × 1,024 / 9,216 = 697.8. Measured with --random-range-ratio 1.0, --warmup-requests 64 and --flush-cache, so no request reuses another's prefix.
8× MI350X (1×8, prefill and decode unified) cookbook Balanced cell, random 8,192 in / 1,024 out, concurrency 16 · SGLang v0.5.19 @ 12771786 (ROCm), TP8, fp8 KV
462 tok/s per accelerator
Marked Final Verification In Progress in the cookbook: P50 TTFT 6,331 ms, P50 TPOT 32.8 ms and 462 tok/s per GPU at concurrency 16. The MI350X and MI355X share one recipe: TP8 on ROCm with AITER's A8W4 MoE kernels, Triton attention, an fp8_e4m3 KV cache and, in this cell, no speculation. The per-GPU figure counts prompt and generated tokens: 16 × 1,024 / (6.331 + 1,024 × 0.0328) = 410.4 generated tok/s, against 462 × 8 × 1,024 / 9,216 = 410.7. Not linked to a catalog system, which carries the MI355X and not the MI350X.
GB300 (TP8) single user (batch 1), 8K in / 1K out · vLLM day-0 build, FP8 KV cache
111 tok/s per user while decoding
Per-user decode speed at batch 1, not aggregate serving throughput — the two differ by orders of magnitude and must not be compared directly. With DSpark speculative decoding the same configuration reaches 331 tokens/s per user, a 3.14x speed-up. Linked to the GB300 NVL72 with the 8 GPUs the run used. CORRECTED 2026-09-13: this note said the row was not linked because the catalog has no 8-GPU GB300 system. A run on part of a rack links the rack and records the GPUs it used, as the SGLang rows beside it do.
GB300 (TP16) single user (batch 1), 8K in / 1K out · vLLM day-0 build, FP8 KV cache
118 tok/s per user while decoding
Per-user decode speed at batch 1 across 16 accelerators; 370 tokens/s with DSpark speculative decoding. Doubling tensor parallelism from TP8 buys only 6% single-user speed, which is the point: latency scales poorly with GPU count while aggregate throughput scales well. Linked to the GB300 NVL72 with the 16 GPUs the run used.

Kimi K3

16× GB300: 1 PP8 prefill worker → 1 TP8 decode node (1P:1D) serving frontier at its throughput end, 8K in / 1K out, concurrency 1024 · SGLang + Miles, prefill/decode disaggregated, fp4 arm
2,808 tok/s per accelerator · ≈ 44,928 tok/s across 16× GB300
The throughput end of the serving frontier in SGLang's Kimi K3 post: 2,808 tok/s per GPU at 18.7 tok/s per user, with a client concurrency limit of 1,024 printed beside the point. The figure counts prompt and generated tokens and divides by all 16 GPUs of the arm, the prefill worker's included, so the run served about 44,928 tok/s, about 4,992 of them generated; at 18.7 tok/s per user that is about 267 streams decoding at once, well inside the 1,024 limit. The post does not state that denominator; the post's source record sets out the measurements it is inferred from. The arm is tagged fp4, which the post never defines, and the figure says its points come from a separate sweep; the untagged PP8 → TP8 arm peaks at 2,715. That sweep's concurrency-1 point decodes 115.8 tok/s per user, beside the ~113 tok/s the post gives for batch 1 before speculation and far from the ~423 it gives with DSpark, so the sweep most likely ran without speculation; the post does not say. CORRECTED 2026-09-13: this row was sglang-k3-gb300-8gpu-batched and stored 22,464 tok/s, 2,808 decode tokens/s per GPU times 8 accelerators, left unlinked as an 8-GPU topology. The figure counts prompt tokens as well, and the run used 16 GPUs.
16× GB300: 1 PP8 prefill worker → 1 TP8 decode node (1P:1D) serving frontier, untagged PP8 → TP8 arm at its peak, 8K in / 1K out · SGLang + Miles, prefill/decode disaggregated
2,715 tok/s per accelerator · ≈ 43,440 tok/s across 16× GB300
The PP8 → TP8 · 1P:1D arm of the same frontier at its highest point, 2,715.1 tok/s per GPU at 18.74 tok/s per user, read from the SVG's data coordinates, whose axis ticks fit a straight line exactly; the figure prints no concurrency for it. It is the fp4 arm's topology without the fp4 tag. 16 GPUs × 2,715 = 43,440 tok/s counting prompt and generated tokens, about 4,827 of them generated, or about 258 streams at 18.74 tok/s. The per-GPU figure is divided by all 16 GPUs, the prefill worker's included: the post does not say so, and its source record infers it from the post's own measurements. The post names no precision and no speculation setting for its untagged arms. The caption calls arms within about 2% of one another indistinguishable.
32× GB300: 2 PP8 prefill workers → 2 DCP8 decode nodes (2P:2D) serving frontier, the DCP composition at its peak, 8K in / 1K out · SGLang + Miles, prefill/decode disaggregated, TP8 + DCP8 decode
2,633 tok/s per accelerator · ≈ 84,256 tok/s across 32× GB300
The caption's DCP composition, two PP8 prefill workers feeding two DCP8 decode nodes, at 2,633 tok/s per GPU; the figure plots it at 2,632.7 and 18.44 tok/s per user and prints no concurrency. Decode context parallelism shards the MLA cache inside each TP8 group, so a DCP8 decode node is still 8 GPUs and the arm is 32. 32 GPUs × 2,633 = 84,256 tok/s counting prompt and generated tokens, about 9,362 of them generated, or about 508 streams at 18.44 tok/s. The per-GPU figure is divided by all 32 GPUs, both prefill workers included: the post does not say so, and its source record infers it from the post's own measurements. The post names no precision and no speculation setting for its untagged arms. A single PP8 → DCP8 · 1P:1D arm peaks at 2,682.7 on 16 GPUs, within the 2% the caption calls indistinguishable.
8× GB300 (2×4 trays), TP8 + DCP8 + host-memory KV tier AgentX replay of coding-agent sessions, 48 concurrent sessions · SGLang + Miles, DCP8, hicache ratio 2
aggregate 541 tok/s · ≈ 68 tok/s per accelerator (8× GB300)
The post's caption gives 541 tok/s at 48 sessions without saying what it counts, and only the alt text of the figure beside it calls that rate aggregate decode throughput. The row files it as unspecified, because the figure itself fits neither reading, as below. The run replays real coding-agent sessions on 2×4 GB300, 8 GPUs, with a host-memory KV tier at ratio 2 on both arms and decode context parallelism as the only difference; the TP8 arm collapses at 16 sessions, from 3,925 to 1,410 tok/s per GPU. The post names no precision and no speculation setting for the replay. The figure beside that text does not plot 541. Its axis counts prompt and generated tokens per GPU, and the DCP8 arm's 48-session point reads about 15,333 tok/s per GPU, about 122,660 across the 8 GPUs, at 8.9 tok/s per stream; 48 streams at 8.9 tok/s generate about 427 tok/s. 541 is neither that nor 122,660, and the post reconciles it with neither. The DCP8 curve peaks at 48 sessions and falls to about 12,750 at 56 and 8,520 at 64. Read from the SVG's data coordinates; the DCP8 panel has no axis labels of its own and is read on the TP8 panel's scale. CORRECTED 2026-09-13: this row said 541 was roughly 2.4% of the same hardware's short-context batched figure. That set a generated-token rate for 8 GPUs against 22,464, a prompt-inclusive per-GPU figure multiplied by 8, and the row linked no hardware although it ran on GB300.
Inference Economics · static, local-first calculationCatalog 0.8.1 · cutoff 2026-09-13 · app v0.8.1 · build 1c759ca