HomeArticle

CPU is back in the game, kicking off a $170 billion power play

36氪的朋友们2026-06-22 09:35
CPU is the most unexpected variable in the current AI cycle. As AI evolves from conversations to Agents, the demand for CPU in inference has surpassed that in training.

On June 1st, NVIDIA launched the Vera CPU at the GTC Taipei 2026 conference during the Computex Taipei. The newly launched next - generation AI supercomputing platform, Vera Rubin, has its first - batch customers including OpenAI and Anthropic.

This is NVIDIA's first time to launch an independent CPU product line. NVIDIA's growth in the past 20 years has been almost entirely based on GPUs. NVIDIA CEO Jensen Huang said at the press conference that in the era of AI agents, the CPU has become a key bottleneck in data - center performance. We can't let the CPU slow down the token production speed of the AI factory.

In May, AMD CEO Lisa Su announced in the earnings conference call that the market - size forecast for server CPUs has been doubled from $60 billion to over $120 billion, and the compound annual growth rate from 2025 to 2030 has been raised from 18% to 35%.

According to IDC statistics, the global server market reached $444.1 billion in 2025, a year - on - year increase of 80.4%, with AI servers contributing most of the increment. UBS predicted in a recent semiconductor - industry research report that the potential market size of server CPUs will grow from about $30 billion in 2025 to about $170 billion in 2030, nearly a five - fold increase in five years.

Data from market research firm Mercury Research shows that in the first quarter of 2026, AMD's server - CPU revenue share reached 46.2%, while Intel's was 53.8%. However, AMD's shipment share was only 33.2%, and Intel still accounted for 66.8%. That is to say, AMD created higher revenue with fewer chips, and the premium ability of high - core - count products was prominently reflected in this quarter.

Lin Meibing, the chief analyst at ICTIME, told Economic Observer that the CPU is the most unexpected variable in the current AI cycle. As AI evolves from dialogue to agents, the demand for CPUs in inference has exceeded that in training.

01

The GPU is "waiting" for the CPU

In November 2025, Intel and the Georgia Institute of Technology jointly published a paper titled "A CPU - Centric Perspective on Agentic AI". In this paper, the research team conducted actual tests on five typical agent workloads. The results showed that the time occupied by CPU - side tool processing accounted for 43.8% to 90.6% of the total latency.

A securities analyst who has long tracked the semiconductor sector said that during the large - model training phase, the CPU's workload accounts for only about 10% to 30%, and in some workloads, it may reach nearly 40%. Most of the calculations are borne by the GPU. This is because the calculation process of AI large - model training is highly regular. Hundreds of millions of parameters perform matrix multiplications repeatedly on massive data. The GPU's parallel architecture is designed for such tasks, and the CPU is responsible for data loading, communication scheduling, and result copying, without involving core matrix operations.

However, in the inference phase, this ratio begins to reverse. The CPU's workload share rises to over 70% and is even higher in agent scenarios. Because agent tasks require multi - step inference, calling external tools, executing code, reading and writing databases, searching the web, and then arranging the intermediate results into the final output.

Programming assistants, data - analysis tools, and automated research agents all belong to this category and are currently the fastest - growing scenarios in large - model applications. The common feature of these tasks is that they are control - flow intensive, have complex branches, and frequent input and output. The utilization rate of the GPU will significantly decrease when facing such serial and fragmented tasks.

Many industry insiders said that in agent tasks, the overall utilization rate of the GPU is generally less than 50%, far lower than the 70% to 85% of traditional inference services. The token consumption of AI deployment in the agent mode is usually 20 to 30 times that of ordinary conversations because a single user interaction often involves dozens of tool calls and intermediate inferences.

According to IDC's prediction, the annual number of global agent - executed tasks will grow from about 44 billion in 2025 to over 400 trillion in 2030.

Intel's management said in the first - quarter 2026 earnings conference call that in the era of AI agents, the number of CPU cores required per gigawatt of power consumption may grow from the current approximately 30 million to 120 million. Market research firm Gartner also predicts that by 2027, 40% of agent projects will be scaled back or cancelled due to infrastructure - cost overruns, and a significant portion of these overruns comes from the continuous tool - call and context - management overhead on the CPU side.

Agents generate a large amount of intermediate data when processing long conversations and complex tasks. The AI system needs to remember all previous conversation contents and tool - call results during the inference process, which is called KV Cache (key - value cache) in industry terms. It will expand continuously with the number of conversation rounds, but the built - in storage capacity of the GPU is very limited. NVIDIA's H100 only has 80GB, and the next - generation B200 only has 192GB. The intermediate data generated by a complex agent task can easily exceed this limit.

Currently, the common approach in the industry is to transfer this intermediate data from the GPU to the CPU side. The CPU can be externally connected to DDR5 memory, with a single - chip capacity reaching several terabytes, one to two orders of magnitude larger than the GPU storage.

In November 2025, the CXL industry alliance, composed of chip manufacturers such as Intel, AMD, and ARM, released the CXL 4.0 protocol (Compute Express Link, an open standard for high - speed interconnection between chips), allowing multiple CPUs to share the same large - capacity memory pool and reducing the overhead of data transfer between chips.

As a result, the CPU is no longer only responsible for task scheduling but also for data storage and memory management during the AI inference process.

Li Bin, the vice - president of R & D at Beijing Qingwei Intelligence Technology Co., Ltd., said that when the AI computing - power cluster expands to the scale of a super - node with thousands of chips directly connected, the complexity of interconnection and scheduling between chips increases exponentially. Qingwei Intelligence's solution can already achieve direct connection of 4096 chips, and the interconnection cost is significantly reduced compared with traditional solutions.

Li Bin said that the super - node technology itself is not new. It is the growth of the parameter scale of large models that has made this kind of architecture useful. In a cluster of this scale, the coordination and management workload borne by the CPU far exceeds that of small - scale deployments.

In addition, the CPU itself has also undergone intensive technological upgrades in the past few years. The core count of server CPUs has climbed from 28 cores in 2017 to 288 cores (Intel Clearwater Forest) and 256 cores (AMD Venice) in 2026, with a nearly ten - fold increase in density.

In 2023, Intel introduced the AMX (Advanced Matrix Extension) instruction set, enabling the CPU to have a dedicated matrix - computing unit for the first time. According to Intel's test data, in the deep - learning inference scenario, the AI performance of the fourth - generation Xeon processor equipped with AMX is nearly ten times higher than that of the previous generation at most. The memory subsystem has also been upgraded from DDR4 to DDR5, with both the bandwidth and capacity of a single platform doubling.

The upgrade of the core count and instruction set also corresponds to the change in the CPU - to - GPU ratio. Intel CEO Patrick Gelsinger said in the first - quarter 2026 earnings conference call that in the training scenario, usually 7 to 8 GPUs are paired with 1 CPU, in the inference scenario, it converges to 3 to 4 GPUs per CPU, and in the agent scenario, it is expected to further converge to 1:1.

Intel CFO David Zinsner added in the same conference call that the overall CPU - to - GPU ratio in the industry has converged from 1:8 in the past to about 1:4.

02

The first major price increase in over a decade

The above - mentioned ratio change has been reflected in product pricing.

Jia Bin, the market - department head of a CPU distributor in Shenzhen, told reporters that since February 2026, Intel and AMD have successively raised the prices of their entire series of server CPUs, with an overall increase of 10% to 15%. The spot premium of some high - end AI server CPUs is even higher, and there may be a new round of price increases in the second half of the year.

Jia Bin said that in the past more than a decade, server CPUs basically "increased in quantity without increasing in price". The performance improved with the process upgrade, but the unit price remained unchanged. This year's price increase is rare in the industry. The capacity utilization rate of Intel's main production line has risen from less than 80% to 100%, and multiple models are out of stock, with a delivery cycle of 3 to 4 months.

AMD is also facing tight capacity. Jia Bin said that 2026 is the first time since he entered the industry that he has seen the server - CPU capacities of Intel and AMD almost fully booked. "In the past, the CPU supply was always sufficient, but this year it's the opposite."

Jia Bin also noticed that customers' demand for CPUs when purchasing AI servers is being divided into two categories. One is the CPU that cooperates with GPU operations inside the cabinet, which pursues the ultimate core count, over 128 cores, with an average price of over $4000, while the average price of traditional server CPUs is only over $2000. The other is the CPU independently deployed outside the cabinet, used for tool execution, sandbox operation, and task arrangement of agents. It doesn't require ultimate performance, about 64 cores are enough, but the quantity required is much larger.

Jia Bin said that in an ideal state, each agent task occupies one CPU exclusively, and independent deployment is more efficient than virtualized partitioning. The average price of CPUs outside the cabinet is about $3000. "The higher the core count, the greater the increase in the unit price, not a proportional increase. So, it is a common practice for customers to use mid - range products outside the cabinet to increase the quantity and flagship products inside the cabinet to ensure performance."

In a semiconductor - industry research report titled "Rise of the Agents" released on June 11th, Bank of America Securities raised its forecast for the total potential market size (TAM) of server CPUs in 2030 to over $170 billion and for the first time split this market into three parts: about $30 billion for traditional cloud - computing CPUs, about $70 billion for AI - cluster head - node CPUs, and about $70 billion for AI - agent independent - node CPUs. Among them, the size of the third part was close to zero in 2025 and is a brand - new market that emerged in 2026.

Morgan Stanley also predicted in a research report on June 4th that agent - based AI will bring an additional demand of $32.5 billion to $60 billion to the server - CPU market before 2030. Zhongtai Securities defined 2026 as "the first year for CPUs to benefit from the AI boom" in a in - depth CPU research report released on June 7th.

The above - mentioned Bank of America Securities research report also listed a set of historical shipment - volume comparisons: In 2022, the shipment volume of AI CPUs was equivalent to 19% of the shipment volume of AI accelerators (such as GPUs). By 2025, this ratio rose to 51%, and it is expected to reach 127% in 2030. According to this prediction, the number of CPUs in AI servers will exceed that of GPUs within five years.

03

New demands for domestic CPUs

Information released by NVIDIA during the Computex Taipei shows that its newly launched Vera CPU is based on the ARM architecture (a CPU instruction set known for low power consumption and high energy efficiency, paralleled with x86 as the two major mainstream architectures). Up to 256 units can be deployed in a single cabinet, and it uses liquid - cooling for heat dissipation.

In the agent - sandbox scenario, Vera's performance is 1.8 times that of x86 processors. In NVIDIA's newly launched Vera Rubin supercomputing cluster (NVIDIA's next - generation AI data - center platform), a 40 - rack POD (the smallest complete computing unit composed of multiple racks) contains 1152 Rubin GPUs and up to 1088 Vera CPUs, with a ratio close to 1:1.

NVIDIA also mentioned that the previously launched Grace CPU has shipped nearly 2.5 million units in total, and the CPU - related revenue in 2026 is expected to be close to $20 billion.

Jia Bin believes that the statistical caliber of the above - mentioned $20 billion is relatively broad, covering the revenue attribution of CPUs in various product forms, which is not exactly the same as the revenue from selling CPU chips separately in the traditional sense. However, even considering the caliber difference, this is already a considerable volume for a company that did not have an independent CPU business in 2024.

Lin Meibing believes that the symbolic significance of NVIDIA's entry into the CPU market is greater than the product itself. In the past, AI servers were centered around GPUs, and CPUs were just supporting components. When the world's largest GPU company enters the CPU market and locks in OpenAI and Anthropic as its first - batch customers, the market position of the CPU is completely different from what it was two years ago.

According to AMD's first - quarter 2026 earnings report, the company's data - center business revenue reached $5.775 billion, exceeding Intel's $5.1 billion in the same period for the first time. Moreover, Lisa Su proposed a five - year goal in the earnings conference call: to achieve an annual data - center revenue of $100 billion.

Intel CEO Patrick Gelsinger also said on multiple public occasions that he has firm confidence in the core role of the CPU in the AI era.

This is also an opportunity for Chinese CPU - industry chain enterprises. Jia Bin said that domestic leading cloud providers are increasing their procurement of server CPUs this year. On the one hand, it is to purchase CPUs to match GPUs for newly built AI data centers. On the other hand, because the CPU - to - GPU ratio has converged from 1:8 in the past to 1:4 or even higher, the number of CPUs required for the same data center is more than twice that of last year.

In fact, a relatively complete industry chain has been formed around server CPUs in China.

Haiguang Information (688041.SH) is one of the domestic manufacturers with the largest shipment volume of x86 - architecture server CPUs. According to relevant earnings reports, Haiguang Information's revenue in 2025 was 14.377 billion yuan, a year - on - year increase of 56.92%. In the first quarter of 2026, the revenue was 4.034 billion yuan, and the year - on - year growth rate further increased to 68.06%.

Ying Zhiwei, the vice - president of Haiguang Information, told reporters that the competition in AI computing power has shifted from the performance comparison of single chips to system - level collaboration. Haiguang is one of the few domestic chip - design companies with both high - end general - purpose CPU and DCU (AI accelerator) product lines. The collaboration between the two product lines can meet the full - scenario requirements of AI training and inference. The Haiguang C86 series of CPUs (compatible with the x86 instruction set) have also built a domestic open ecosystem with more than 6000 partners.

Ying Zhiwei said that more than 90% of the existing applications in the commercial field are developed based on the x86 architecture. This ecological compatibility is Haiguang's core advantage in the replacement of domestic products for information - technology innovation. In addition, in addition to data centers, edge inference and embodied intelligence are also becoming new growth scenarios for x86 CPUs. For example, the mainstream solutions of robot manufacturers such as Ubtech and Unitree still mainly use x86 CPUs paired with GPUs, and the software ecosystem of x86 has an obvious first - mover advantage in the field of embodied intelligence.

According to public information, Huawei's Kunpeng follows the ARM full - stack self - developed route. The Kunpeng 920/950 is deeply coordinated with Ascend AI chips and mainly serves Huawei's own ecosystem and the information - technology innovation market.

In terms of supporting chips, the main product of Montage Technology (Shanghai) Group Co., Ltd. (688008.SH) is the memory - interface chip (a signal - transfer chip between the server CPU and the memory module). According to public information, its memory - interface chips ranked first in the global market with a 36.8%