NVIDIA · accelerator · 80 GB · — official claim
1 × H100 SXM · 0.7 kW · researched
to cutoff 2026-07-19
Unpriced accelerator; use an explicit scenario acquisition price.
A local inference system only wins under a workload, a utilization rate, and a performance profile. Explore the crossover point against hosted API spend — with every assumption in view.
Adjust the workload and operating envelope. Results recompute locally and are shareable as a URL.
Every value is versioned, dated, and tagged with the kind of evidence behind it.
Measured observations stay separate from modeled profiles. Follow the source registry for context.
27,417.3 tokens / sec · MLPerf Inference v5.0 · server
measured high confidenceServer result; benchmark harness and latency constraints apply.
30,724.5 tokens / sec · MLPerf Inference v5.0 · server
measured high confidenceServer result; benchmark harness and latency constraints apply.
92,338.7 tokens / sec · MLPerf Inference v5.0 preview · server
measured medium confidencePreview result; benchmark harness and latency constraints apply.
57,177 tokens / sec · MLPerf Inference v4.1 · server
measured high confidence50,796 tokens / sec · MLPerf Inference v4.1 · server
measured high confidence22,021 tokens / sec · AMD ROCm · latency-constrained comparison
measured medium confidenceAMD page reports 22,021 tok/s versus H100 21,605; CPU, runtime, and latency caveats apply.
21,605 tokens / sec · AMD ROCm · latency-constrained comparison
measured medium confidenceAMD page comparison; CPU, runtime, and latency caveats apply.
24,110 tokens / sec · AMD ROCm · latency-constrained comparison
measured medium confidenceAMD page reports 24,110 tok/s versus H100 24,525; CPU, runtime, and latency caveats apply.
24,525 tokens / sec · AMD ROCm · latency-constrained comparison
measured medium confidenceAMD page comparison; CPU, runtime, and latency caveats apply.
NVIDIA · accessed 2026-07-19
Open source ↗NVIDIA Newsroom · accessed 2026-07-19
Open source ↗Mercatus AI · accessed 2026-07-19
Open source ↗NVIDIA · accessed 2026-07-19
Open source ↗NVIDIA · accessed 2026-07-19
Open source ↗Reuters (Investing.com reprint) · accessed 2026-07-19
Open source ↗NVIDIA · accessed 2026-07-19
Open source ↗NVIDIA · accessed 2026-07-19
Open source ↗AMD · accessed 2026-07-19
Open source ↗AMD · accessed 2026-07-19
Open source ↗AMD · accessed 2026-07-19
Open source ↗Moonshot AI · accessed 2026-07-19
Open source ↗Moonshot AI · accessed 2026-07-19
Open source ↗DeepSeek · accessed 2026-07-19
Open source ↗Qwen · accessed 2026-07-19
Open source ↗Meta · accessed 2026-07-19
Open source ↗Mistral AI · accessed 2026-07-19
Open source ↗Mistral AI · accessed 2026-07-19
Open source ↗Mistral AI · accessed 2026-07-19
Open source ↗Anthropic · accessed 2026-07-19
Open source ↗OpenAI · accessed 2026-07-19
Open source ↗OpenAI · accessed 2026-07-19
Open source ↗OpenAI · accessed 2026-07-19
Open source ↗Google · accessed 2026-07-19
Open source ↗xAI · accessed 2026-07-19
Open source ↗Mistral AI · accessed 2026-07-19
Open source ↗MLCommons · accessed 2026-07-19
Open source ↗NVIDIA · accessed 2026-07-19
Open source ↗AMD ROCm · accessed 2026-07-19
Open source ↗Reproducible formulas, explicit exclusions, and a clear quality caveat.
Per-request input, cache-hit, cache-write, and output tokens are priced separately. Throughput is constrained by both decode and prefill; facility electricity applies IT power × PUE. Support is an annual ratio of all-in acquisition cost. Break-even exists only when hourly API-equivalent savings exceed self-hosted variable cost.
The quality-parity multiplier scales the value of successful local output against the API comparator. It does not claim the models are equivalent, and it never changes physical throughput.
Financing, taxes, staffing, downtime, resale, opportunity cost, and production support
beyond the configured ratio are excluded. The research cutoff is the catalog's last-updated
boundary; changing the URL scenario changes assumptions, not the evidence. The repository's docs/architecture.md and docs/PROJECT.md record the full scope.