HomeArticle

HUAWEI's Chip and AI Game

36氪的朋友们2026-10-08 13:41
If in the past the competition in the AI industry first revolved around who could develop more powerful models, then shifted to who could provide more computing power, the next stage where participants can truly widen the gap from others lies in the capability to integrate models, computing, energy, capital and the ecosystem into a continuously evolving system. That is exactly Huawei's real game.

Recently, Huawei has intensively released the views of Xu Zhijun and Yu Chengdong to the public. Xu Zhijun emphasized that Huawei has formed its own technical route in multi-chip interconnection and system architecture. At the same time, Yu Chengdong also stressed that Huawei's access channels to advanced semiconductor processes are limited, but chips such as Ascend and Kirin have achieved significant improvements in energy efficiency and overall system-level performance. The views of both of them represent that the basic unit of current AI competition is expanding from single-point technology to systems, ecosystems and infrastructure systems.

In the past few years, the most striking part of artificial intelligence competition has always been models. From GPT to DeepSeek, from parameter scale, benchmark tests to reasoning capabilities, every leap of the model affects people's judgment on the AI competition pattern. However, as the model scale continues to expand and the reasoning complexity keeps increasing, the factors restricting the further expansion of AI are extending to the underlying infrastructure: the availability of computing power, chip interconnection, data transmission, large-scale cluster collaboration, and the energy supply of data centers have all begun to directly affect the scale and efficiency at which model capabilities can enter the real world.

Data from the International Energy Agency (IEA) in 2026 provides an intuitive scale: the capital expenditure of five large technology companies in 2025 has exceeded 400 billion US dollars, and it is expected to increase by about 75% in 2026; the capacity of cutting-edge data centers built specifically for AI has more than tripled in the past 18 months. AI is increasingly showing industrial characteristics of capital intensity, energy intensity and long-cycle infrastructure intensity. It is still driven by algorithm innovation, but it can no longer rely solely on algorithm expansion. This change brings a more important consequence: the basic unit of AI competition is changing.

01 From Chips to Systems: The Basic Unit of AI Competition is Moving Upward

In the past, AI hardware competition mainly focused on a single GPU, with process technology, single-card computing power, memory capacity and bandwidth forming the core indicators. These indicators are still important to this day, but with the continuous expansion of computing scale, the performance of a single chip has become increasingly difficult to represent the actual computing power of large-scale AI systems. When thousands, tens of thousands or even more accelerators jointly perform training and reasoning tasks, what the model finally obtains is not the nominal peak computing power of a certain chip, but the effective computing power that the entire system can continuously deliver. Data exchange between chips, memory access, network topology, task scheduling, fault recovery, as well as power supply and heat dissipation, all begin to affect the conversion efficiency of theoretical computing power to effective computing power. The basic unit of AI competition has also expanded step by step: from chips, computing nodes and super nodes, further extending to computing clusters, data centers and even artificial intelligence infrastructure. Chips still determine the physical starting point of computing capabilities, but as more and more computing units are incorporated into the same system, the final performance increasingly depends on the combined effect of chip performance and system efficiency.

Herbert Simon pointed out when discussing complex systems that such systems are usually hierarchical. A complex system is composed of subsystems at different levels, and its overall behavior depends not only on the attributes of the components, but also on the organization and interaction between these parts. When AI computing enters the stage of large-scale systems, competitive advantages will inevitably expand from the performance of a single device to the collaborative capability between different computing levels.

It is in this change that Huawei's original technical accumulation has gained new strategic significance. If the competition mainly stays at the level of a single chip, Huawei is still facing a difficult catch-up: factors such as advanced process, high-bandwidth memory, manufacturing capacity and production capacity will directly limit the room for improvement of its single-chip performance. However, as the competition scale expands from single chips to large-scale computing systems, Huawei has gained another space to improve the overall computing power: improving the collaboration efficiency of a large number of chips through high-speed interconnection and system architecture innovation.

At the Huawei Connect 2026, Huawei clearly shifted its strategic focus to AI infrastructure, and proposed a large-scale computing route with "super node + cluster" as the core. It seems that Huawei is focusing on "how many cards are connected", but in fact it is trying to redefine the problem of computing competition: in addition to continuing to improve the performance of a single chip, can a large number of computing units be organized into a more efficient whole through system architecture innovation? The so-called system architecture innovation includes several levels: the super node architecture that organizes a large number of Ascend chips into larger computing units; high-speed interconnection that reduces the data exchange bottleneck between chips through UnifiedBus, all-optical interconnection, etc.; the cluster architecture that further organizes multiple super nodes into large-scale computing clusters; plus supporting software scheduling, storage, fault tolerance and energy management. In other words, the "system architecture" here covers the entire computing organization above the chip.

02 How System Advantages are Transformed into Ecological Advantages: Huawei's Next Threshold

If Huawei is only understood as an AI chip company catching up with NVIDIA, its special position in this competition will be underestimated. What Huawei has in hand is not an isolated chip technology, but a set of capabilities that were distributed in different industries in the past and now recombined by AI: Ascend computing, Kunpeng general computing, communication networks, optical interconnection, storage, servers, data centers, digital energy, and long-accumulated system engineering capabilities. AI is pulling these technologies into the same system: large models require accelerators, accelerators require high-speed interconnection, large-scale clusters require storage, scheduling and fault tolerance, and when the power demand of data centers continues to rise, power supply and distribution and cooling also begin to directly affect the cost and expansion speed of computing power. The technology portfolio formed by Huawei across the ICT industry in the past has gained new strategic value and has become different parts of the same AI infrastructure.

Huawei uses "black soil" as a metaphor for the basic technology platform, hoping that Kunpeng, Ascend, basic software and cloud infrastructure provide the underlying environment for partners to develop models and industry applications on top of it. The really important part of this metaphor is that it reveals a deeper logic of infrastructure: the value of infrastructure is not only reflected in its own performance, but also depends on the scale of external innovation it can carry, and whether these innovations can in turn enhance the entire technology system. Ecosystem theory summarizes this relationship as "complementarity": the platform is interdependent with the products, services and innovations built on it, and the expansion of participants may further increase the value of the entire ecosystem. In this sense, the real test of Huawei's so-called "black soil" is not only what performance Ascend itself can achieve, but also whether it can form a technical foundation sufficient to attract model companies, developers and application enterprises to continue to participate.

This touches on Huawei's most challenging card - the software and developer ecosystem. The advantage that NVIDIA is hardest to replicate is never just a GPU. After long-term accumulation, CUDA has connected development tools, computing libraries, performance optimization, engineering experience, talent cultivation and existing code into an interdependent technology ecosystem, thus forming significant migration costs.

This advantage has typical characteristics of increasing returns and path dependence. Brian Arthur's research on competitive technologies points out that when a technology becomes more valuable as users increase, experience accumulates and complementary resources expand, the advantages formed in the early stage may continuously self-reinforce through positive feedback and form obvious path dependence. The development of CUDA presents a similar mechanism: the expansion of the ecosystem scale continuously enhances the value of the platform, and the improvement of the platform value further attracts developers and complementary resources to enter. Therefore, what Huawei faces is not just the software function gap between CANN and CUDA, but also a technology ecosystem that has been accumulated for a long time and has self-reinforcing effects.

In contrast, Ascend/CANN is still in the process of rapid improvement. In some scenarios, migrating CUDA models to the Ascend platform still requires third-party library adaptation, operator adjustment or development, as well as manual debugging and performance optimization. Huawei is lowering the threshold through CANN open source, compatibility with mainstream frameworks and expanding the developer community, and pushing the goal from "available" to "easy to use", but its actual maturity ultimately needs to be tested in combination with the number of third-party developers, migration costs, native development scale, completeness of development tools, and long-term performance under real workloads.

An important feature of mature infrastructure is that it can hide complexity from users: turn problems solved by engineers into software, precipitate expert experience into tools, and turn one-time optimization into public capabilities that can be called repeatedly. It is also at this node that DeepSeek enters Huawei's game.

03 DeepSeek Moving Down: Models Begin to Participate in Defining Infrastructure

At the end of September 2026, the cooperation between DeepSeek and Huawei further entered the underlying programming infrastructure from model operation. It was reported that DeepSeek is cooperating with Huawei to develop and open source a series of underlying software tools for Ascend, aiming to reduce the technical threshold for developers to use and optimize Ascend chips. This means that the cooperation between the two sides is no longer limited to making DeepSeek models run on Ascend, but has begun to enter the software infrastructure that supports the entire developer ecosystem.

If this is only understood as "domestic models adapting to domestic chips", its industrial significance will be underestimated. Traditional adaptation is still a one-way relationship: the hardware platform already exists, and model manufacturers modify the code and optimize performance to make the model run on a specific platform. The change that has emerged now is that model manufacturers have begun to participate in more underlying links such as computing libraries, communication libraries and programming tools, moving from infrastructure users to co-constructors.

This cooperation can be further understood from the "innovational complementarities" proposed by Timothy Bresnahan and Manuel Trajtenberg in 1995. When studying general purpose technologies, they pointed out that two-way innovation feedback may be formed between basic technologies and downstream applications: the improvement of basic technologies creates new possibilities for application innovation, and the development of downstream applications will generate new technical demands, increasing the benefits of further improving basic technologies. The cooperation between DeepSeek and Huawei presents a similar interactive relationship: the development of models constantly puts forward new computing, communication and software demands, which promote the further optimization of Ascend and its software tools; the improvement of underlying computing capabilities in turn expands the technical space for model training and reasoning.

This kind of innovation complementarity is particularly important in the AI field, because what AI chips ultimately process is not the abstract "artificial intelligence", but the constantly changing specific workloads. Model architecture, operator structure, computing-communication ratio in training, and memory access methods in reasoning will all put forward different requirements for software and hardware, and model companies are exactly the closest to the changes in these demands. Therefore, the significance of DeepSeek's participation in the construction of underlying software tools lies in shortening the feedback distance between model innovation and computing infrastructure, so that the constantly changing demands on the model side can enter the optimization process of software, chips, interconnection and systems faster.

For Huawei, DeepSeek provides workload knowledge from the cutting-edge model side; for DeepSeek, going deep into underlying tools means that it has the opportunity to participate in shaping the computing environment on which the model runs. Whether the two sides can truly form a long-term co-evolutionary ecosystem remains to be observed, but this cooperation has shown that the meaning of "full stack" is changing: it does not necessarily require a company to own all the ownership of chips, clouds, models and applications, and the more critical point is whether different levels of the technology stack can form fast and continuous feedback.

Almost at the same time, AMD provides an interesting reference from the opposite direction. On September 28, 2026, AMD announced that it would acquire World Labs led by Li Feifei in an all-stock transaction of about 8.2 billion US dollars. AMD clearly stated in its official announcement that as AI expands to reasoning, robotics, simulation and physical AI, the workload faced by computing infrastructure will be more diverse, and the model research capability of World Labs can help AMD understand how these workloads evolve, and shape the future hardware, software and system routes accordingly. Lisa Su summarized this capability as: building the next generation of AI computing platforms requires an in-depth understanding of how models evolve.

DeepSeek is moving down into computing infrastructure, while AMD is moving up into model research. The directions are opposite, but they reveal the same structural change: the distance between models and computing infrastructure is shortening. The past relatively clear linear industrial chain of "chip - server - cloud - model - application" is gradually becoming a feedback system: models define workloads, workloads affect software and hardware, and new hardware capabilities change which models are engineering and economically feasible.

04 The Infrastructure of AI Competition and the Possibility of Another Computing System

This feedback system is still extending downward. As computing infrastructure continues to expand, AI competition further touches the energy level. The demand of large-scale data centers for electricity continues to grow, and energy is gradually changing from a background condition supporting computing operation to a pre-factor restricting the construction scale and expansion speed of AI infrastructure. When the scale of AI data centers continues to expand, electricity gradually changes from a background input to a pre-constraint. IEA data shows that global data center electricity consumption increased by 17% in 2025, among which AI-dedicated data centers grew faster; at the same time, the capital expenditure of a few large technology companies has reached a magnitude comparable to traditional energy investment. "AI is becoming a heavy industry" vividly summarizes this trend. More accurately, cutting-edge AI is showing increasingly obvious infrastructure-intensive characteristics, and its development is increasingly dependent on large-scale, long-cycle capital and energy investment. Models can iterate within a few months, chip design and manufacturing are counted in years, data center construction takes longer, and new power generation, grid expansion and transmission facilities are constrained by longer investment and approval cycles. Different technical levels thus present significant time-scale differences: the upper-layer models evolve rapidly, while the underlying infrastructure is difficult to expand synchronously, and a new structural tension is forming between the two.

As land, grid access, energy supply, semiconductor supply chains, data governance and long-term capital increasingly directly affect the construction and expansion of AI capabilities, infrastructure competition has inevitably become an important part of national industrial capability and geopolitical competition. This is also the realistic background for the rise of the concept of "sovereign AI". But sovereign AI is not equivalent to technological self-sufficiency. For the vast majority of economies, it is neither realistic nor necessarily economical to achieve full-chain autonomy from semiconductor equipment, advanced chips to clouds, models and applications. A more accurate understanding is to reduce single-point dependence on key links and retain necessary strategic options: a country still participates in the global technical division of labor, while hoping that key computing power, data and model capabilities will not be completely interrupted due to changes in a single supplier or external policies. Future AI infrastructure will not necessarily evolve into several technical worlds completely isolated from each other, but will more likely form interdependence of different degrees and structures. In this sense, the competition unit has ushered in another expansion: from enterprise competition to ecological competition, and further to infrastructure system competition jointly shaped by enterprises, supply chains, energy systems, capital and national policies.

The change of competition scale also makes Huawei's significance go beyond the technical competition between single enterprises. What Huawei is promoting is not only the research and development of domestic chips, but also testing whether another set of computing systems can be formed under external technical constraints. This route has dual significance: on the one hand, Huawei tries to use its existing advantages in communications, optical networks and system engineering to open up new space for performance improvement at the system level; on the other hand, it also provides an alternative path to reduce the single-point dependence of key AI infrastructure on specific external technologies and suppliers. But this strategic value is not equivalent to technological leadership, nor can it be deduced that a relatively complete computing system built around chips, interconnection, software and data centers will inevitably be formed. It ultimately has to accept the market test of performance, cost, reliability and developer ecosystem.

05 Huawei's Cards Also Include the Cards It Lacks

Looking at "Huawei's game" from this perspective, its advantages and constraints are quite clear. The first card is still Ascend. Without its own AI accelerator, the rest of the system capabilities will lack the computing core; the second card is interconnection and system engineering. Huawei is advancing the accumulated communication and optical network capabilities from "connecting devices" to "organizing computing", and UnifiedBus, NPO and super nodes all serve to improve the collaboration efficiency of large-scale computing units; the third card is a relatively complete ICT infrastructure portfolio. Kunpeng, storage, networks, data centers and digital energy enable Huawei to design AI systems on a larger scale than a single chip; the fourth card is the emerging software and model ecosystem, including CANN opening and model developers such as DeepSeek beginning to participate in the construction of underlying tools. The fourth card is currently relatively weak, but it may determine how much value the first three can ultimately exert, because the first three can certainly be promoted by Huawei's own R&D investment, but the ecosystem cannot be created by a single enterprise alone. Whether the ecosystem is established ultimately depends on whether enough external developers and enterprises are willing to enter and can stay independently.

The cards Huawei lacks are also obvious. First of all, advanced manufacturing capabilities and stable production capacity. System architecture can improve the utilization efficiency of existing chips, but it cannot eliminate manufacturing constraints. Huawei's AI computing devices still face the problem that demand exceeds supply capacity. Secondly, the time accumulation of the software ecosystem. The libraries, tools, talents and code assets formed by CUDA over many years have obvious path dependence, which is difficult to replicate through short-term investment. Third, the engineering cost of large-scale systems themselves: using more chips to make up for the gap of single chips means higher pressure on communication, power supply, heat dissipation, fault and operation and maintenance. Therefore, "scale" is not a free substitute for advanced chips, but an engineering exchange. Finally, Huawei's competitors are also moving in the same direction. NVIDIA, AMD and large cloud vendors are all expanding from single devices or cloud services to a more complete AI infrastructure. Therefore, system competition is not Huawei's unique route, but more like a common shift happening in the entire AI industry.

This also makes the criteria for judging Huawei's route more stringent. System innovation can indeed change the weight of various technical capabilities in competition, but it cannot eliminate physical constraints; it is more like a