Series โ Accelerator Measured Benchmarks
Owner: Finance / Charlie AGT-002 Source ID: SRC-AI-MLCOMMONS (MLPerf Training/Inference), vendor reproducible disclosures Unit: benchmark-native (time-to-train, tokens/s, samples/s); perf-per-chip and perf-per-watt derived where system size is disclosed Cadence: per MLPerf round (~2 per year) + event-driven vendor disclosures Visibility: PUBLIC As-of convention: every row carries the MLPerf round or disclosure date and a source grade; derived perf/$ uses list or disclosed pricing with the price source noted
How to append: prefer MLCommons closed-division results (Primary). Vendor self-reported numbers outside MLPerf are ๐ก at best and must be labeled. Never compare across precisions or divisions without noting it.
| As-of | System (chip ร count) | Benchmark / round | Result | Per-chip derived | Perf/W or Perf/$ (note basis) | Source | Grade |
|---|---|---|---|---|---|---|---|
| 2025-11-12 | NVIDIA Blackwell x 5,120 (GB200/GB300 NVL72 clusters) | MLPerf Training v5.1 โ Llama 3.1 405B pretraining | 10 min time-to-train (round record) | n/a | n/a | https://mlcommons.org/2025/11/training-v5-1-results/ + https://developer.nvidia.com/blog/nvidia-blackwell-architecture-sweeps-mlperf-training-v5-1-benchmarks/ | ๐ก (figure via vendor blog; confirm against MLCommons closed-division results table) |
| 2025-11-12 | NVIDIA Blackwell x 2,560 | MLPerf Training v5.1 โ Llama 3.1 405B pretraining | 18.79 min (45% faster than prior-round 2,496-GPU submission) | n/a | n/a | https://developer.nvidia.com/blog/nvidia-blackwell-architecture-sweeps-mlperf-training-v5-1-benchmarks/ | ๐ก (vendor blog; MLPerf round is Primary) |
| 2025-11-12 | NVIDIA GB300 NVL72 (72x Blackwell Ultra; Lambda submission) | MLPerf Training v5.1 | 27% faster than GB200 NVL72 at same scale (per Lambda; see source for benchmark detail) | n/a | n/a | https://lambda.ai/blog/lambda-mlperf-training-benchmarks-v5.1 | ๐ก (submitter blog) |
| 2025-11-12 | NVIDIA GB300 NVL72 vs Hopper (same GPU count) | MLPerf Training v5.1 โ Llama 3.1 405B pretrain / Llama 2 70B LoRA | >4x Hopper (405B pretrain), ~5x (70B LoRA) at equal GPU count | generation-over-generation per-GPU ratio | n/a | https://blogs.nvidia.com/blog/mlperf-training-benchmark-blackwell-ultra/ | ๐ก (vendor blog; cross-generation ratio, not absolute) |
| 2025-09-10 | NVIDIA GB300 NVL72 (Blackwell Ultra MLPerf debut) | MLPerf Inference v5.1 (published 2025-09) | New inference records in Blackwell Ultra debut round incl. DeepSeek-R1 and Llama 3.1 405B categories (figures in source) | n/a | n/a | https://developer.nvidia.com/blog/nvidia-blackwell-ultra-sets-new-inference-records-in-mlperf-debut/ + https://www.hpcwire.com/2025/09/10/mlperf-inference-v5-1-results-land-with-new-benchmarks-and-record-participation/ | ๐ก (vendor summary of Primary round) |