//Inference Economics research cutoff · 2026-07-19
Open reference · v0.1

Price the decision, expose the uncertainty.

A local inference system only wins under a workload, a utilization rate, and a performance profile. Explore the crossover point against hosted API spend — with every assumption in view.

01 · model a scenario

Calculator

Adjust the workload and operating envelope. Results recompute locally and are shareable as a URL.

Live result

02 · inspect the inputs

Catalog

Every value is versioned, dated, and tagged with the kind of evidence behind it.

H100 SXM
NVIDIA · accelerator · 80 GB · — official claim
1 × H100 SXM · 0.7 kW · researched to cutoff 2026-07-19
Unpriced accelerator; use an explicit scenario acquisition price.
H200 8-GPU system
NVIDIA · node · 1,128 GB · $370,000 assumption
8 × H200 · 10.2 kW · researched to cutoff 2026-07-19
price range $320,000–$420,000 ·Independent typical 8-GPU estimate; planning power assumption.
DGX B200
NVIDIA · node · 1,440 GB · — official claim
8 × B200 · 14.3 kW · researched to cutoff 2026-07-19
Full-node acquisition price not published; planning assumption required.
DGX B300
NVIDIA · node · 2,304 GB · $550,000 estimated
8 × B300 · 14.5 kW · researched to cutoff 2026-07-19
Independent US price report, with 1.20 deployment planning multiplier.
GB200 NVL72
NVIDIA · rack · 13,400 GB · — assumption
72 × GB200 · 132 kW · researched to cutoff 2026-07-19
Rack price not published; planning assumption required.
GB300 NVL72
NVIDIA · rack · 20,736 GB · $3,000,000 assumption
72 × GB300 · 132 kW · researched to cutoff 2026-07-19
Rack price and power are planning assumptions; 3.0m is not MSRP.
Instinct MI300X 8-GPU platform
AMD · node · 1,536 GB · — assumption
8 × MI300X · 8.5 kW · researched to cutoff 2026-07-19
Full-node power planning value; acquisition price not published.
Instinct MI325X 8-GPU platform
AMD · node · 2,048 GB · — assumption
8 × MI325X · 10.5 kW · researched to cutoff 2026-07-19
Full-node power planning value; acquisition price not published.
Instinct MI355X 8-GPU platform
AMD · node · 2,304 GB · $325,000 assumption
8 × MI355X · 14 kW · researched to cutoff 2026-07-19
Node acquisition and 14 kW planning assumptions; not MSRP.
03 · calibrate your confidence

Evidence

Measured observations stay separate from modeled profiles. Follow the source registry for context.

Benchmark observations

8×H200 · Llama 2 70B

27,417.3 tokens / sec · MLPerf Inference v5.0 · server

measured high confidence

Server result; benchmark harness and latency constraints apply.

8×MI325X · Llama 2 70B

30,724.5 tokens / sec · MLPerf Inference v5.0 · server

measured high confidence

Server result; benchmark harness and latency constraints apply.

DGX B200 · Llama 2 70B

92,338.7 tokens / sec · MLPerf Inference v5.0 preview · server

measured medium confidence

Preview result; benchmark harness and latency constraints apply.

8×H200 · Mixtral 8x7B

57,177 tokens / sec · MLPerf Inference v4.1 · server

measured high confidence
8×H100 · Mixtral 8x7B

50,796 tokens / sec · MLPerf Inference v4.1 · server

measured high confidence
MI300X server · Llama 2 70B

22,021 tokens / sec · AMD ROCm · latency-constrained comparison

measured medium confidence

AMD page reports 22,021 tok/s versus H100 21,605; CPU, runtime, and latency caveats apply.

H100 server · Llama 2 70B

21,605 tokens / sec · AMD ROCm · latency-constrained comparison

measured medium confidence

AMD page comparison; CPU, runtime, and latency caveats apply.

MI300X server · Llama 3.1 405B

24,110 tokens / sec · AMD ROCm · latency-constrained comparison

measured medium confidence

AMD page reports 24,110 tok/s versus H100 24,525; CPU, runtime, and latency caveats apply.

H100 server · Llama 3.1 405B

24,525 tokens / sec · AMD ROCm · latency-constrained comparison

measured medium confidence

AMD page comparison; CPU, runtime, and latency caveats apply.

Source registry

NVIDIA Supercharges Hopper: H200

NVIDIA Newsroom · accessed 2026-07-19

Open source ↗
Exclusive: prices of Nvidia B300 server

Reuters (Investing.com reprint) · accessed 2026-07-19

Open source ↗
AMD Instinct MI300X Platform Data Sheet

AMD · accessed 2026-07-19

Open source ↗
AMD Instinct MI325X Platform Datasheet

AMD · accessed 2026-07-19

Open source ↗
Mistral Large 3 model card

Mistral AI · accessed 2026-07-19

Open source ↗
Mistral Large 3 675B Instruct repository

Mistral AI · accessed 2026-07-19

Open source ↗
MLPerf Inference v5.0 results

MLCommons · accessed 2026-07-19

Open source ↗
MLPerf Inference v4.1 results

NVIDIA · accessed 2026-07-19

Open source ↗
04 · read the frame

Methodology

Reproducible formulas, explicit exclusions, and a clear quality caveat.

What is calculated

Per-request input, cache-hit, cache-write, and output tokens are priced separately. Throughput is constrained by both decode and prefill; facility electricity applies IT power × PUE. Support is an annual ratio of all-in acquisition cost. Break-even exists only when hourly API-equivalent savings exceed self-hosted variable cost.

Quality-parity caveat

The quality-parity multiplier scales the value of successful local output against the API comparator. It does not claim the models are equivalent, and it never changes physical throughput.

Exclusions and semantics

Financing, taxes, staffing, downtime, resale, opportunity cost, and production support beyond the configured ratio are excluded. The research cutoff is the catalog's last-updated boundary; changing the URL scenario changes assumptions, not the evidence. The repository's docs/architecture.md and docs/PROJECT.md record the full scope.

Inference Economics · static, local-first calculationCatalog version 0.1.0 · research cutoff 2026-07-19