HomeArticle

Ahead of the AMD AI Conference, NVIDIA showcased its impressive CPU achievements, the Vera Rubin has entered full mass production, and more than 300 partners have completed relevant deployment.

36氪的朋友们2026-07-22 09:08
NVIDIA is taking the battle from the GPU market to the CPU market.

NVIDIA is expanding its battle from the GPU market to the CPU arena.

On the eve of AMD's AI-related event, NVIDIA unveiled the latest progress of its next-generation Vera Rubin platform on Tuesday, the 21st Eastern U.S. Time: the Vera Rubin NVL72 has entered the ramp-up phase of mass production. NVIDIA states that the platform's supply chain spans more than 350 factories across 30 countries, with over 300 partners involved.

At the same time, NVIDIA further released performance data of the Vera CPU and Vera Rubin platform in AI agent workloads, attempting to prove that as AI evolves from "answering questions" to autonomously planning, invoking tools, and executing tasks, CPUs are becoming a new battlefield for AI infrastructure.

The full specifications, benchmark test results, and architecture information of NVIDIA's disclosed data center CPU product Vera are critical data essential for potential customers to comprehensively evaluate the chip. The company stated that Vera chips were delivered to customers including OpenAI, Anthropic, and SpaceX as early as June.

The launch of Vera marks the latest step in NVIDIA's vertical integration strategy. The company intends to sell customers complete rack systems containing self-developed chips, rather than simply selling individual chips. According to Wolfe Research's estimates, the average price of a single Vera chip is approximately $5,000, with projected shipments reaching about 1.3 million units this year. For AMD and Intel, NVIDIA's move directly threatens the core interests of both companies in the server CPU market.

This is not the first time NVIDIA has announced that the Vera CPU has entered mass production. At the end of May this year, NVIDIA announced that Vera had entered full-scale production, claiming that the chip could complete specific tasks 1.8 times faster than traditional x86 CPUs. In mid-May, NVIDIA also delivered the first batch of Vera CPU systems to customers including Anthropic, OpenAI, and SpaceX. The focus of this announcement has further shifted to the mass production ramp-up, customer deployment, and real-world production environment performance of the entire Vera Rubin platform.

01

The Rise of AI Agents Makes CPUs a "Critical Bottleneck" Again

The core logic behind NVIDIA's bet on CPUs is that AI applications are undergoing fundamental changes.

The main task of traditional generative AI is to generate answers based on user input, with GPUs handling large-scale model computations while CPUs are mostly responsible for relatively peripheral scheduling tasks. However, with the rapid development of AI agents, AI systems have begun to independently break down tasks, invoke external tools, run code, access data, and repeatedly evaluate results.

In this process, CPUs need to frequently process a large number of low-latency, real-time tasks. NVIDIA believes that AI agents are not workloads that rely solely on GPUs: every agent's runtime environment, tool invocation, task orchestration, and long-context data retrieval all require CPU participation.

When NVIDIA previously released Vera, it pointed out that agentic AI is creating a new "CPU moment." Its judgment is that as AI systems shift from "answering questions" to "taking actions," the role of CPUs in the entire AI factory will significantly increase.

NVIDIA even predicts that the long-term scale of the server CPU market could reach $200 billion. Compared with the traditional server CPU market, this figure appears extremely aggressive, but NVIDIA's logic is that the future market boundaries of CPUs may no longer be limited to traditional enterprise servers, but extend to AI inference, AI agents, reinforcement learning, data processing, and various control and orchestration tasks for AI infrastructure.

This is also why NVIDIA is attempting to redefine the rules of CPU competition.

02

Not Competing on Core Count, NVIDIA Bets on "Single-Core Speed"

Vera's design path differs significantly from traditional server CPUs.

NVIDIA claims that Vera is its first CPU specifically designed for agentic AI, featuring 88 self-designed Olympus cores and 1.2TB/s of memory bandwidth. NVIDIA previously stated that Vera can achieve a 50% single-core performance improvement over traditional CPUs; in tests announced at the end of May, Vera completed specific tasks 1.8 times faster than x86 CPUs.

In its latest disclosure, NVIDIA further emphasized Vera's core design orientation: rather than simply stacking more cores, Vera prioritizes single-threaded performance, inter-core communication bandwidth, and memory access latency.

NVIDIA states that Vera's custom Olympus cores can deliver 2x single-threaded performance, 3x inter-core bandwidth, and 40% lower memory latency compared to competing chiplet designs. The goal is to enable AI agents to complete tasks faster and return computing resources to the GPU as soon as possible, thereby improving the utilization of the entire AI factory.

The underlying business logic is very straightforward: GPUs are one of the most expensive and scarce computing resources in AI data centers. If the CPU processes tasks too slowly, the GPU may be left in a waiting state. NVIDIA hopes to reduce this "idling" time by specifically optimizing the CPU.

In production environment tests of the Vera CPU conducted by DeepInfra, NVIDIA claims that Vera can support up to 1.6 times more concurrent AI agents and achieve up to 2.2 times faster task orchestration. It is important to note that these figures come from customers' specific test environments and do not mean that Vera comprehensively outperforms all general-purpose CPU workloads.

03

The Competitive Focus Has Escalated From a Single Chip to an Entire "AI Factory"

The Vera CPU is just one part of NVIDIA's larger system strategy.

In the Vera Rubin platform, NVIDIA not only integrates its self-developed Vera CPU, Rubin GPU, and network and data processing chips, but also incorporates the Groq 3 LPX inference acceleration rack built on Groq technology, forming a complete AI infrastructure covering training, post-training, test-time scaling, and real-time inference.

NVIDIA states that the Vera Rubin NVL72 is ramping up mass production globally, with partners including CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure already deploying related racks. NVIDIA also notes that its supply chain covers more than 350 factories across 30 countries, with over 300 participating partners.

In terms of customer testing, CoreWeave reported after running DeepSeek-R1 tests on the Vera Rubin NVL72 that compared to the previous-generation Grace Blackwell NVL72, the number of Tokens generated per megawatt per second increased by 10 times. Google Cloud has launched A5X instances based on the Vera Rubin NVL72, which NVIDIA claims can deliver lower inference costs and higher Token throughput per megawatt in specific scenarios.

This means NVIDIA's competitors are no longer just individual accelerator chips like AMD's Instinct GPUs or Intel's Gaudi series.

What NVIDIA is trying to control is the entire infrastructure: CPUs handle scheduling, GPUs handle computation, network chips handle high-speed interconnection, software handles scheduling and optimization, and the final product is delivered to customers as a rack or even a data center-level system.

For AMD and Intel, the threat of this competitive approach is that NVIDIA may not need to compete with them in a fully symmetrical manner in the traditional CPU market. NVIDIA's goal is to cut into the most growth-potential parts of AI servers, first seizing workloads such as agentic AI, reinforcement learning, and high-performance AI inference, and then gradually expanding the application scope of CPUs.

04

AMD and Intel's Traditional Advantages Face New Challenges

For a long time, Intel and AMD have held dominant positions in the server CPU market.

But NVIDIA's current judgment is that AI infrastructure is changing the value hierarchy of CPUs. In the past, server CPU competition revolved more around core count, general-purpose computing capabilities, and overall throughput; in the era of agentic AI, the importance of low latency, single-threaded performance, memory bandwidth, and task orchestration efficiency is on the rise.

This is exactly the market gap that NVIDIA believes it can step into.

In terms of product form, Vera can be used as a standalone CPU or combined with NVIDIA GPUs to form a Vera Rubin system. NVIDIA also positions it as a key component connecting AI agents and GPU computing resources.

NVIDIA's previously announced customer list already includes AI companies such as Anthropic, OpenAI, and SpaceX, as well as cloud service providers like ByteDance, CoreWeave, and Oracle Cloud Infrastructure; server vendors including Dell, Hewlett Packard Enterprise, Lenovo, and Supermicro are also advancing systems based on Vera.

This gives NVIDIA an advantage different from traditional CPU vendors: it can provide customers with a complete solution through the combination of GPUs, CPUs, networks, and software.

But this does not mean NVIDIA can easily shake AMD and Intel.

Competition in the CPU market is not just about chip performance, but also includes software ecosystems, customer certifications, server compatibility, and long-term supply relationships. Especially for large cloud service providers and enterprise customers, it usually takes years to complete the transition of hardware platforms.

Therefore, whether Vera can truly expand from AI-specific scenarios to the broader server market still depends on whether customers can obtain sufficiently obvious performance and cost advantages.

05

NVIDIA Is Unlocking a New Revenue Curve Beyond GPUs

NVIDIA's concentrated release of Vera Rubin's mass production and performance information on the eve of AMD's AI event also carries obvious market competitive implications.

In the past few years, NVIDIA has almost exclusively occupied the most core profit pool of AI infrastructure relying on GPUs. But as AMD continues to expand its AI accelerator layout and Intel also tries to find a breakthrough in the AI chip market, the market is increasingly focusing on whether NVIDIA can maintain its rapid growth.

At the same time, the CPU businesses of AMD and Intel have regained market attention due to the new demand expectations brought by agentic AI.

NVIDIA's response is: if the value of AI data centers is expanding from single GPU computing to complete AI factories, then NVIDIA does not have to surrender the CPU market to competitors.

From this perspective, Vera is not just a server CPU simply launched by NVIDIA, but a further extension of its vertical integration strategy.

NVIDIA hopes that customers will no longer buy individual GPUs, CPUs, or network chips, but a complete set of infrastructure that can directly run AI models and AI agents.

If agentic AI truly becomes the core driving force for the next round of computing demand growth, then what NVIDIA will compete for may not just be GPU market share, but the entire AI server computing value chain.

This also means that AMD and Intel's future competitors may no longer just be each other, but an NVIDIA that already has GPU, CPU, network, software, and complete system capabilities.

This article is from the WeChat Official Account "Wall Street News Max", written by LI Dan & YANG Chen, and published with authorization from 36Kr.