NVDA, AMD · US
Vera Rubin NVL72 vs GB200: inference TCO & architecture | SemiAnalysis note
SemiAnalysis: Rubin inference tokens/MW & TCO. https://newsletter.semianalysis.com/p/vera-rubin-nvl72-vs-gb200-nvl72-inference
Cite a section with a deep link, e.g. /en/r/sa-vera-rubin-nvl72-tco-2026#thesis
Snapshot
- Source date
- 2026-07-23
- Key metric
- tok/s per MW & TCO
- Early lead (note)
- ~5× perf/$ class
- Stage
- Early bring-up
Structured research note summarizing a SemiAnalysis subscriber article (not a full reprint). Defer to the original for detail.
Thesis
Using early CoreWeave silicon plus SemiAnalysis InferenceX/TCO models: Rubin NVL72 leads GB200 on inference tokens/MW and $/token, but is still in bring-up—software maturity should widen the gap. Read-through is next-gen rack economics reinforcing NVDA system premium, not a raw-FLOPS contest.
Analysis
Frame
The note stress-tests Nvidia claims vs GB200 NVL72 (incl. early-2025 bring-up), InferenceX Jul-2026 baselines, and public MI355X distributed inference. Metrics shift from peak FLOPS to tokens/s per MW and perf per TCO (AI TCO / Datacenter models).
Takeaways (note-level)
- Early samples: Rubin on DeepSeek R1 cited around ~5.4× perf/MW and ~5× perf/$ vs GB200, with runway as software matures
- CoreWeave public messaging emphasized much higher tokens/MW at matched interactivity—cross-check methodology
- At high interactivity, older platforms struggle to serve; rack co-design is the wedge
Mapping
NVIDIA system pricing power; see NVDA and GPU utilization. AMD remains disadvantaged on rack-scale inference unless software/world-size catch up.
Implications
| Theme | Read-through |
|---|---|
| Inference economics | $/token + MW > marketed FLOPS |
| Rack generation | NVL72 co-design moat |
Outlook
Watch Rubin production software (Dynamo/TRT-LLM), more third-party benches, GB300 comps, share in power-constrained regions.
Risks
References
- Original (SemiAnalysis): Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis
- InferenceX mirror of the note — https://inferencex.semianalysis.com/blog/vera-rubin-nvl72-vs-gb200-nvl72-inference
- CoreWeave: Vera Rubin tokens per MW blog — https://www.coreweave.com/blog/nvidia-vera-rubin-nvl72-on-coreweave-10x-more-tokens-per-megawatt-than-blackwell
- NVIDIA Vera Rubin blog hub — https://blogs.nvidia.com/blog/vera-rubin/
- SemiAnalysis archive — https://newsletter.semianalysis.com/archive
Not investment advice. Copyright remains with SemiAnalysis / authors; this package is an on-site research index.