HomeArticle

Advance Screening AMD Advancing AI 2026: AMD AI YES?

AI唱反调2026-07-22 17:28
What is still missing from AMD's AI factory?

The AMD Advancing AI 2026 conference kicks off today, with Lisa Su delivering her keynote speech tomorrow (July 23). The event is expected to unveil the Zen 6 EPYC Venice processor and the MI455X series AI accelerator. The MI500, part of AMD's roadmap targeting a 1000x increase in AI performance within four years and slated for a 2027 release, will not be officially launched tomorrow.

AI clusters are "construction sites for building large models", where the focus is on stacking computing cards and competing for training speed, with NVIDIA as the absolute dominant player. AI factories, on the other hand, are "production lines for applying large models", where the competition centers on inference cost and deployment efficiency. This track has not yet formed a monopoly, and AMD believes it has a real opportunity.

However, the concept of "AI factory" is not an original AMD idea. NVIDIA launched its DGX AI Factory solution back in 2025, and has already secured orders from the vast majority of the world's supercomputing and cloud vendors. AMD's stated shift "from cluster to factory" essentially implies that it cannot outperform NVIDIA in the cluster track and is looking to claim a share of the factory market. NVIDIA defines a closed factory using its proprietary NVLink protocol, while AMD counters this control over standards by leveraging open specifications including Helios, UALink and UEC. The stickiness of a closed ecosystem is clearly far stronger.

When it comes to AI, can AMD finally deliver a resounding "YES" this time?

What products will be launched at tomorrow's conference

AMD CTO Mark Papermaster has already given a preview: the Zen 6 architecture will make its official debut tomorrow, marking a full generational architecture upgrade.

The Zen 6 Venice is manufactured using TSMC's 2nm process. The standard configuration offers a maximum of 96 cores, while the density-optimized Zen 6c variant supports up to 256 cores and 512 threads. AMD claims it delivers over 70% performance improvement compared to the Zen 5 Turin platform. It features the new SP7 socket, a TDP range of 700-1400W, 16-channel DDR5 memory, PCIe 6.0 support, and a memory bandwidth of 1.6TB/s per socket.

On the GPU side, the spotlight will most likely fall on the MI455X. The product was previously previewed at CES 2026: it is built on the CDNA5 architecture, equipped with 432GB of HBM4 memory, and delivers 40 PFLOPS of FP4 performance. When paired with the Venice CPU, it forms the Helios rack-level reference design, with 4 MI455X accelerators and 1 Venice CPU per node, delivering 2.9 exaFLOPS of FP4 performance per full rack. Meta has already announced that it will deploy Helios-based racks in the second half of 2026, and OpenAI has signed a multi-generation cooperation order for 6GW of MI450 products.

However, one detail in the benchmark results deserves attention: AMD claims that at the 100kW rack level, Venice delivers 3.3 times the performance of NVIDIA Vera. This 3.3x figure is derived from model-based estimation rather than actual physical testing. AMD explicitly notes in its methodology document: "These results are intended to provide directional comparisons, not direct measurements of rack-level benchmarks."

So that 3.3x number should be taken with a grain of salt.

AMD's "small pie" in the inference market is indeed well-baked

All the orders AMD has secured in the inference market are tangible, revenue-generating contracts. What's more, AMD has adopted a unique "equity-for-orders" mechanism for these deals.

Meta now runs 100% of its Llama 405B inference workloads on MI300X hardware, under a multi-year contract valued at $60-10 billion. The deal includes up to 160 million AMD warrants, with a share price unlock threshold set at $600. OpenAI has also signed a 6GW MI450 cooperation agreement, which similarly includes 160 million warrants. A single tranche of warrants covers a maximum of 160 million shares, accounting for roughly 10% of AMD's total share capital. Combined, the two deals could dilute existing shareholders by up to 20% if all warrants are exercised — but this is conditional on the phased completion of 6GW deployment between 2026 and 2030, and AMD's share price reaching $600.

This "equity-for-orders" model essentially turns Meta and OpenAI into "quasi-shareholders", which is a far stronger countermeasure against NVIDIA's ecosystem lock-in than competing on pure cost-performance alone.

Of course, this is a double-edged sword: if the share price fails to meet the exercise threshold over the long term, the attractiveness of the warrants will diminish. If the warrants are successfully exercised, the roughly 20% equity dilution will also reduce returns for existing shareholders.

The MI300X is priced very competitively. On cloud platforms, it is priced at $1.80-2.30 per GPU-hour, 40-60% cheaper than the $5.01 per GPU-hour for the H100. It is equipped with 192GB of HBM3 memory, more than double the 80GB of the H100, eliminating the need to split large model inference workloads across multiple cards. In MLPerf Inference benchmarks, the MI355X delivers 92-104% of the performance of NVIDIA's B300, meaning the performance gap has become extremely narrow.

ROCm 7.x also offers far better support for mainstream open-source frameworks including PyTorch, vLLM, Llama and Qwen than it did two years ago. CUDA's moat is becoming shallower — while it has not been completely filled in, AMD is no longer the outsider that "can't even install PyTorch".

More critically, Microsoft Azure officially announced on July 20 that it will begin large-scale deployment of Helios-based racks, and has added two new EPYC Venice-powered virtual machine series: HDv2 optimized for agentic AI and data pipelines, and HXv2 optimized for EDA workloads. Satya Nadella explicitly stated that this offering will provide customers with "enhanced performance, greater scale, and expanded choice". Microsoft is AMD's first full-stack hyperscale cloud customer covering GPU, CPU and networking. Helios itself is an open standard reference design, manufactured and branded by ODMs including Sanmina, Wiwynn, Wistron and Inventec in compliance with ORW specifications. AMD's customer base has now expanded from "hyperscalers with long-term contracts" to "cloud vendors selling computing resources", marking a genuine second step forward in its business model evolution.

All these figures are clearly documented in AMD's Q1 earnings report: data center revenue reached $5.8 billion, a 57% year-over-year increase. Q2 total revenue guidance is set at $11.2 billion, representing a 46% year-over-year rise, with the data center business expected to post an even higher year-over-year growth rate in Q2.

AMD has indeed baked a very solid small pie in the inference market. But a complete "AI factory" requires not only the downstream baking workshop, but also an upstream "flour mill".

How many more steps does AMD need to take to build its AI factory?

A true "AI factory" requires a fully closed loop covering both training and inference. Training sits upstream of the workflow, while inference sits downstream. Whoever controls the training market gets to define model architectures, framework standards and developer habits.

NVIDIA's data center revenue for FY2026 hit $193.7 billion, while AMD's data center revenue for FY2025 was approximately $16.6 billion — a full order of magnitude difference. Comparing the first quarters of FY2026, NVIDIA recorded roughly $39.1 billion in revenue versus AMD's $5.8 billion, leaving a 6.7x gap. NVIDIA holds over 80% of the training market share, and this figure will not shift in the short term.

When it comes to real-world training performance, the MI300X's actual results are far less impressive. While its theoretical FLOPS figures look impressive, third-party testing shows it only delivers less than 30% of its theoretical peak performance on real training workloads, compared to 40% for the H100. The gap widens further in multi-node training scenarios, where the H100 outperforms it by 10-25%, with the difference growing even larger at larger scales. The conclusion from SemiAnalysis is straightforward: AMD's training performance per dollar of cost is inferior to NVIDIA's.

While AMD can win over inference market customers with cost-performance advantages, in the training market, the cost of customer migration far exceeds the price of hardware — it requires full re-learning of the entire software stack, toolchains and developer workflows. This cannot be solved simply with money, and requires significant time investment. And time is on NVIDIA's side.

Worse still, AMD's existing space in the inference market is also facing growing pressure. Domestic Chinese products including the Ascend 950 super-node and the Sugon 8000 fully domestic 100,000-card cluster made their debut at WAIC 2026. Chinese domestic large models are deeply optimized for domestic chips, forming a fully closed loop of "domestic models + domestic chips". The more immediate pressure comes from cost-performance: the per-computing-unit cost of domestic Ascend-series inference chips is 30-50% lower than that of the MI300X, and Chinese government and enterprise projects have explicit requirements for domestic component adoption. However, this competitive pressure is largely limited to the Chinese domestic market. In the scenarios of mainstream overseas cloud vendors and hyperscale customers, domestic computing hardware currently has no competitive edge, so AMD's overseas inference market base remains stable.

AMD's "AI factory" narrative is built on real foundations, but it also faces real shortcomings. The real foundations: top-tier customers including Meta, Microsoft and OpenAI are indeed running inference workloads on AMD silicon, the orders are real, and the 57% revenue growth rate is also a verified fact. Helios is a proven reference design, the Venice processor is already in high-volume production at TSMC's 2nm facility, and four ODMs are prepared for mass manufacturing — though the full revenue realization curve will largely play out in 2027. The real shortcoming: without breaking into the training market, the "AI factory" will lack the upstream "raw material workshop", meaning it can only ever be half a complete factory.

AMD is framing a single well-baked small pie as if it were an entire factory. The pie tastes good, but a full factory needs to produce many more pies. The most critical test tomorrow will be whether Lisa Su can convince the market that AMD's ovens are already full of many more pies in the process of baking.

Tomorrow's keynote speech and the Q2 earnings report on August 4 are two critical milestones that will validate AMD's "AI factory" narrative. What the market truly wants to hear is not new slogans, but hard, concrete numbers: can AMD share a clear mass production timeline and specific customer list for the MI450, rather than just repeating its "2027 target"? Can the Zen 6 be priced aggressively enough to defend AMD's x86 market base against the combined pressure from NVIDIA Vera and Intel Xeon? Will TSMC's 2nm production capacity become a genuine bottleneck for AMD's mass production rollout?

All these questions will depend on the answers Lisa Su provides tomorrow.

This article is sourced from the WeChat Official Account "AI Contrarian", and is republished by 36Kr with official authorization.