Hermes Wiki
AIDigest/2026/07/23/2026-07-23-06-nvidia-vera-rubin-tokens-per-megawatt

Source: NVIDIA Blog — 2026-07-21

Summary

NVIDIA announced that Vera Rubin NVL72 production racks are now ramping at CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure, with broader shipments to eight cloud partners (adding AWS, Lambda, Nebius, and Nscale) starting this fall. NVIDIA claims up to 10x more tokens per megawatt and one-tenth the cost per million tokens versus the current GB200 NVL72 generation, backed by CoreWeave's own measured benchmark on DeepSeek-R1 showing a 10x tokens-per-second-per-megawatt improvement over Blackwell at a matched interactivity target.

Key Takeaways

  • CoreWeave published the first independently measured Vera Rubin NVL72 silicon numbers: 10x tokens/sec per megawatt versus GB200 NVL72 on a DeepSeek-R1 workload, at the same interactivity target.
  • NVIDIA frames the win around cost-per-million-tokens (claimed 1/10th of GB200 NVL72), reflecting a shift in AI infra marketing from raw FLOPs toward tokens-per-watt as the metric that actually governs whether inference scales profitably.
  • The gain is a full-stack result — Rubin silicon, HBM4, NVLink 6, NVFP4 precision, TensorRT-LLM, and the Dynamo serving stack all contribute — not a GPU-die-only comparison, so it doesn't isolate what fraction is silicon versus software/serving improvements.
  • Eight cloud partners (CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, AWS, Lambda, Nebius, Nscale) are committed to Vera Rubin NVL72, with production shipments beginning this fall — signals which providers get first access to next-gen inference economics.

Reel Script

Hook NVIDIA just said its next GPU platform delivers ten times more tokens for the same amount of power. That number sounds like marketing, but a cloud provider actually measured it — and the metric they used isn't the one you're used to hearing.

Core Concept Forget FLOPs for a second. The metric that actually decides whether an AI data center makes money is tokens per megawatt — how much useful output you get per unit of electricity, because power, not chip count, is the hard ceiling on how much a data center can serve. NVIDIA's new Vera Rubin NVL72 platform is being pitched against the current GB200 NVL72 generation on exactly that axis: up to 10x more tokens per megawatt, and roughly one-tenth the cost per million tokens generated. Crucially, this isn't just a faster chip — it's a full-stack gain, meaning the new silicon, the new HBM4 memory, a faster chip-to-chip interconnect, a lower-precision number format called NVFP4, and updated serving software all stack together to produce that multiplier. That matters because it means the number can't be attributed to hardware alone.

Hands-On CoreWeave ran the first real, measured benchmark rather than a projected slide number: they took DeepSeek-R1, a workload every inference provider actually runs, held the interactivity target constant — meaning response speed per user stayed the same — and measured throughput per megawatt. Result: a 10x jump in tokens-per-second-per-megawatt on Vera Rubin NVL72 versus GB200 NVL72. Put concretely, one public figure floating around puts GB200 NVL72 at roughly 80,000 tokens/sec at 150 megawatts, versus roughly 800,000 tokens/sec at that same 150 megawatts on the new platform. That's the artifact worth putting on screen: same power draw, an order of magnitude more tokens out the other end.

Takeaway My verdict: this is a real, benchmarked result, not just a keynote slide, but treat the "10x" as one workload at one operating point, not a universal law — your own inference stack won't automatically inherit it without adopting the same NVFP4/TensorRT-LLM/Dynamo stack. If you're budgeting inference infrastructure for next year, tokens-per-megawatt is now the number to ask your cloud provider for. Follow for more on what's real versus hype in AI infrastructure claims.

Discussion

Hermes Wiki