Nvidia's next generation AI rack (codenamed Vera Rubin) claims 5x the performance per dollar over its current generation, based on one partner's engineering samples rather than an independent benchmark still months away.
Nvidia's next-generation AI datacenter rack, codenamed Vera Rubin NVL72, claims to deliver roughly 5x the performance per dollar and 5.4x the performance per megawatt over the company's current-generation GB200 NVL72 rack (the flagship Blackwell-based system). The headline figure comes from a single partner's engineering samples, not an independent benchmark. The race to verify it starts in the next few months.
That figure leads a new SemiAnalysis analysis of the new rack. It is not a verdict. The 5x and 5.4x numbers come from engineering samples at CoreWeave, a CoreWeave customer (and Nvidia partner) that brought up the first Rubin rack. SemiAnalysis has not independently verified them, and the brief labels them as such. CoreWeave's own blog frames the same measurement as "10x more tokens per megawatt" than Blackwell, a different metric cut on the same underlying data. Until someone outside that vendor-to-customer loop runs the workload, the headline is a starting position.
Vera Rubin NVL72 is the second generation of Nvidia's rack-scale architecture, where 72 GPUs act as a single compute unit connected by a high-bandwidth fabric. The inference gain, per SemiAnalysis's companion architecture piece, comes from "extreme co-design" between silicon, networking, and software. The first public release of the Rubin software stack shipped via Nvidia's CUDA 13.4 developer preview, with upstreamed pull requests to PyTorch, vLLM, and OpenAI's Triton compiler. That is more than a paper launch.
One technical point eases the bring-up. Blackwell could not reuse Hopper's WGMMA kernels (the low-level matrix-multiply instructions tuned for the prior generation), so software had to be rewritten. Rubin can reuse Blackwell kernels, which shrinks the engineering bill for early customers. Reaching peak performance still requires fresh kernel work, and the next step is harder: Feynman, the architecture after Rubin, is a much more complex kernel transition than Blackwell to Rubin. Nvidia has telegraphed a more aggressive software cliff on its own roadmap.
Rubin introduces a 3-bit programmable LUT tensor core, a matrix-math unit that can act as a small lookup table, trading a bit of precision for higher throughput on certain inference paths. It is the kind of feature that buys real numbers on specific workloads and is invisible on others, which is why vendor benchmarks and independent benchmarks tend to diverge.
The independent benchmark in question is InferenceX, a SemiAnalysis-maintained framework with a published traces dataset that anyone can run against. Nvidia has committed to submitting verifiable Rubin numbers to InferenceX by the third quarter of 2026. Google is expected to submit its next-generation TPU (TPUv7) in the next couple of months. AMD has committed MI455X UALoE72 (a competing rack-scale platform) submissions. Until those land, the rack-scale inference race has Nvidia's number on the board and nobody else's.
That is the part to watch. The 5x figure is a snapshot from one rack at one partner on one workload. It is a defensible starting claim, and it might hold up under independent testing. It is also the kind of number that will shape the next wave of AI datacenter capex, and a smart reader should treat it as a starting position, not a finish line. The finish line is the InferenceX leaderboard sometime in the second half of 2026, with Google's and AMD's entries on the same page.