NVDA, AMD · US

Vera Rubin NVL72 vs GB200: inference TCO & architecture | SemiAnalysis note

SemiAnalysis: Rubin inference tokens/MW & TCO. https://newsletter.semianalysis.com/p/vera-rubin-nvl72-vs-gb200-nvl72-inference

Published Updated Open interactive reader

Cite a section with a deep link, e.g. /en/r/sa-vera-rubin-nvl72-tco-2026#thesis

Snapshot

Source date
2026-07-23
Key metric
tok/s per MW & TCO
Early lead (note)
~5× perf/$ class
Stage
Early bring-up
As of 2026-08-03

Structured research note summarizing a SemiAnalysis subscriber article (not a full reprint). Defer to the original for detail.

Thesis

Using early CoreWeave silicon plus SemiAnalysis InferenceX/TCO models: Rubin NVL72 leads GB200 on inference tokens/MW and $/token, but is still in bring-up—software maturity should widen the gap. Read-through is next-gen rack economics reinforcing NVDA system premium, not a raw-FLOPS contest.

Analysis

Frame

The note stress-tests Nvidia claims vs GB200 NVL72 (incl. early-2025 bring-up), InferenceX Jul-2026 baselines, and public MI355X distributed inference. Metrics shift from peak FLOPS to tokens/s per MW and perf per TCO (AI TCO / Datacenter models).

Takeaways (note-level)

  • Early samples: Rubin on DeepSeek R1 cited around ~5.4× perf/MW and ~5× perf/$ vs GB200, with runway as software matures
  • CoreWeave public messaging emphasized much higher tokens/MW at matched interactivity—cross-check methodology
  • At high interactivity, older platforms struggle to serve; rack co-design is the wedge

Mapping

NVIDIA system pricing power; see NVDA and GPU utilization. AMD remains disadvantaged on rack-scale inference unless software/world-size catch up.

Implications

Theme Read-through
Inference economics $/token + MW > marketed FLOPS
Rack generation NVL72 co-design moat

Outlook

Watch Rubin production software (Dynamo/TRT-LLM), more third-party benches, GB300 comps, share in power-constrained regions.

Risks

References

  1. Original (SemiAnalysis): Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis
  2. InferenceX mirror of the note — https://inferencex.semianalysis.com/blog/vera-rubin-nvl72-vs-gb200-nvl72-inference
  3. CoreWeave: Vera Rubin tokens per MW blog — https://www.coreweave.com/blog/nvidia-vera-rubin-nvl72-on-coreweave-10x-more-tokens-per-megawatt-than-blackwell
  4. NVIDIA Vera Rubin blog hub — https://blogs.nvidia.com/blog/vera-rubin/
  5. SemiAnalysis archive — https://newsletter.semianalysis.com/archive

Not investment advice. Copyright remains with SemiAnalysis / authors; this package is an on-site research index.