NVIDIA's Stunning First Benchmark of Vera Rubin: DeepSeek Throughput Surges 30-Fold
This is absolutely mind-blowing.
Just now, NVIDIA publicly released the first on-chip measured data of its next-generation flagship cabinet Vera Rubin NVL72 for the first time in the company's history.
Moreover, this actual test directly deployed DeepSeek-V4-Pro to run the most realistic "agent coding" task!
The results are staggering: compared with the current flagship GB300 NVL72, Vera Rubin's throughput per megawatt has increased by up to 30 times, and the Token cost has dropped by up to 35 times!
Some people exclaimed: "I wouldn't even dare to write such numbers in a PPT to deceive investors!"
Other netizens made a witty comment: "I used to think H200 was expensive at $30,000 per unit, but now looking back, H200 is not even considered a entry-level product anymore."
Let's break it down carefully.
From H200 to GB300, the throughput of real Agent workloads has increased by up to 30 times, while the cost has roughly doubled.
From GB300 to Vera Rubin, the throughput has surged by up to 30 times again, and the price of the entire rack has roughly doubled once more.
According to Jensen Huang's version of Moore's Law, it turns out that NVIDIA's biggest technological breakthrough in the past two years is not the GPU itself, but a 4x price increase that brings a full 900x speed leap!
This time, NVIDIA announced a counterintuitive truth to everyone: the LLM era is over, the Agent era has arrived, and all previous AI benchmark standards are completely invalid!
At the same time, the Vera CPU specially built for agents has been installed by Elon Musk's SpaceXAI overnight, and is even planned to be directly sent into space.
Also today, NVIDIA Groq 3 LPX has entered full mass production.
When Gemma 4 31B runs on it, the speed is absolutely stunning, reaching 3400 Tokens per second directly. The throughput of trillion-parameter models has soared by 35 times.
These new records will completely rewrite the commercial landscape of the Agent ecosystem starting from tonight!
Why did NVIDIA have to launch Vera Rubin?
The latest data from OpenRouter gives the answer: in the real world, an Agent AI task consumes 15 times more Tokens than an ordinary chat conversation!
For example, if you let an Agent research a company to support investment decision-making, during this process, the agent and sub-agents will continuously reason, the accumulated Tokens become the input for the next step, and long-context processing has become the most critical factor limiting Agent AI.
The more agent interactions there are, the higher the throughput — when the number of agents increases by 10 times and tool calls increase by 2 times, the contradiction between computing power and algorithms has spawned new demands.
Don't confuse chatbots with agents, all old benchmarks are obsolete!
At this point, traditional AI benchmark tests (such as fixed-length 8K/1K sequence tests) are completely ineffective.
NVIDIA officials made it clear: performance measurement must evolve! We can no longer test a single inference request, but must capture the complete Agent workflow.
To this end, they used the AgentX benchmark under SemiAnalysis.
This test is no longer a rigid Q&A session, but replays real "coding sessions" with context growth, tool calls and sub-agent generation.
This is why H200 is struggling in the new Agent battlefield, and Vera Rubin is destined to reign supreme.
Vera Rubin NVL72 × DeepSeek-V4-Pro:
The computing power monster is here
To prove the strength of Vera Rubin NVL72, NVIDIA directly used the open-source leader — DeepSeek-V4-Pro (1.6T) for on-chip actual measurement.
The data is out, and the result is astonishing —
Under the AgentX workload, the throughput per megawatt of Vera Rubin NVL72 is a full 30 times higher than that of GB300 NVL72!
Please note that the comparison here is not with the old H200, but with the popular mainstream GB300.
In the same DeepSeek-V4-Pro test, the throughput per megawatt of GB300 is already 15 times higher than that of H200.
However, Vera Rubin has raised it by another 30 times on the basis of GB300, further improving the entire Pareto curve!
What does this mean?
For "AI factories" constrained by power supply, this directly expands the boundary of physical laws.
With the same power budget, Vera Rubin can handle 30 times more agent tasks.
For megawatt-level or even gigawatt-level data centers, this is equivalent to conjuring 30 times more computing power assets out of thin air.
Moreover, NVIDIA DSX MaxLPS technology can perform power management at the GPU, rack and workload levels, allowing up to 40% more GPUs to be configured within the same megawatt budget, further increasing the throughput per megawatt of AI factories!
Token cost drops by 35 times! The "ledger" of the Agent industry has been completely rewritten
Throughput per megawatt directly affects the cost of each generated token. The surge in performance brings the most direct consequence of a nuclear-level explosion in business models.
NVIDIA announced that the cost of Vera Rubin NVL72 to produce per million Tokens is up to 35 times lower than that of GB300 NVL72!
When the inference cost plummets by 35 times, 24/7 digital employees will become a reality, and super applications on the ToC side will see a big boom.
Jensen Huang's move has directly opened the door to the era of Agent applications.
Extreme co-design: How did Jensen Huang achieve this?
You may ask: How can the speed be increased by 30 times? Is it relying on the process of a single chip?
Obviously not.
The amazing performance improvement of Vera Rubin NVL72 relies on "extreme co-design".
Such as disaggregated services, distributed KV cache, KV-aware routing, MegaMoE and so on.
In addition, NVFP4 quantization directly compresses model weights to 4-bit precision, greatly reducing memory usage without sacrificing output quality, allowing throughput to take off instantly.
The 6th generation NVLink and large model MoE provide an interconnection network 10 times faster and 3 times lower latency than off-the-shelf Ethernet, so that large models based on MoE architecture like DeepSeek can smoothly call different "expert" sub-networks between 72 GPUs.
This is no longer just selling graphics cards. What Jensen Huang sells is a complete set of "AI power plants".
NVIDIA Groq 3 LPX is in full mass production:
3400 Token/sec, the era of code agents is changing!
At this year's Hot Chips 2026 conference, NVIDIA also dropped a shocking news: Groq 3 LPX has entered full mass production!
Groq 3 LPX is an exclusive extension for the Vera Rubin NVL72 data center platform.
It is "custom-tailored" for this system, focusing on low-latency inference acceleration.
This black technology was acquired by NVIDIA from startup Groq for a whopping $20 billion last December.