HomeArticle

Why has the computing power competition suddenly shifted to supernodes in the second half of 2026?

数智前线2026-07-27 11:41
Everyone is building super nodes, but their development paths have already diverged.

"Large Chinese internet enterprises will deploy a massive number of super-node solutions in the second half of the year," a senior computing power industry insider told Digital Intelligence Frontline, outlining the market trend. Although the unit price of a super-node server is typically higher than that of a regular server with the same number of accelerators, "the math adds up: it delivers higher computing density and takes up far less physical space. Now that enterprises must pack more computing power into limited floor space, super-nodes are a total 'game-changer'," a H3C representative added.

Even DeepSeek, which had not previously adopted super-nodes, "has recently confirmed it will use the architecture and is currently negotiating pricing," a major cloud computing industry insider told Digital Intelligence Frontline, a move that carries clear industry signaling significance.

This boom was also prominently on display at the 2026 World Artificial Intelligence Conference. Huawei showcased a physical 1024-accelerator super-node made up of 20 cabinets, Sugon demonstrated a 640-accelerator super-node with visible coolant flowing through its immersion cooling system, and Baidu Tianchi, Alibaba Cloud Lingjun Zhenwu M890, Inspur Yingxin MetaBrain, and Muxi Xijing S600 were all placed in prominent positions on their respective booths.

What exactly is a super-node? Simply put, it uses high-speed interconnection technology to deeply integrate a large number of GPUs or accelerators across multiple servers into a logically unified system that acts like a single "supercomputer". This eliminates the bandwidth and latency bottlenecks of cross-machine communication in traditional clusters, drastically improving the utilization rate of GPUs or accelerators and driving down the per-Token cost. "Super-nodes have already become a mainstream trend in current AI Infra development," a senior Huawei official told Digital Intelligence Frontline.

Just a year ago, super-nodes were flagship products used by a handful of vendors to showcase their system capabilities. Today, they are evolving into an arms race that every major computing power enterprise must participate in. "Domestic super-nodes were launched last year, and have now reached the stage where they can be deployed in batches," a Baidu Kunlunxin representative said. "By the end of this year, nearly every computing power enterprise will have the ability to deliver super-node solutions."

Who Is Building Super-Nodes

The super-node concept is no longer new. In 2024, NVIDIA released the GB200 NVL72, which integrated 72 Blackwell GPUs into a single high-speed interconnection domain, bringing the architecture to widespread industry attention, though initial demand was relatively muted. It was not until 2025, as large model parameter counts exploded, that super-nodes became a hot topic. Domestic vendors began technical pre-research work as early as around 2024, with the first batch of products hitting the market in 2025, as Huawei, Baidu Intelligent Cloud, Inspur Information, H3C, ZTE and other players all launched their own offerings. The super-node track began to welcome a flood of new entrants.

However, these vendors are competing on different focal points, and three distinct development paths have seemingly emerged: one camp is continuously expanding the scale of individual super-nodes, another prioritizes deployment efficiency and cost-effectiveness, while the third is attempting to redefine super-nodes using new chip and interconnection architectures.

Digital Intelligence Frontline's observations show that Huawei, Sugon, and Baidu are all moving toward the "large-node" direction.

Huawei has already publicly demonstrated a physical 1024-accelerator Ascend 950 super-node composed of 20 cabinets, and will begin mass delivery of an 8192-accelerator super-node in the fourth quarter of this year — the theoretical maximum full configuration limit for this platform. Huawei's assessment is that as large model parameter counts, context window lengths, and inference concurrency volumes continue to rise, individual computing domains will need to accommodate far more chips and larger memory capacity. The 384-accelerator super-node it launched last year has already seen commercial deployment of more than 750 units across industries including the internet, finance, energy, education, and healthcare, proving super-nodes' capability for large-scale commercial rollout.

Sugon has opted for immersion liquid cooling to support even higher density deployment, with 640 accelerators per cabinet and a single cooling unit supporting two full cabinets. The 100,000-accelerator computing cluster Sugon 8000 (Dengfeng) built on this super-node architecture has already been deployed at the Zhengzhou core node of the National Supercomputing Internet. "It is currently running at 90% capacity, with most demand coming from model training, as well as AI for Science workloads," a Sugon representative told Digital Intelligence Frontline.

Baidu Intelligent Cloud is evolving its Tianchi super-node from a 256-accelerator configuration to a 512-accelerator version, which is scheduled for launch by the end of the year. The platform emphasizes full-stack in-house development spanning interconnection protocols, switching chips, and liquid cooling systems. Du Hai, General Manager of Baidu Intelligent Cloud Hybrid Cloud Division, noted that its super-nodes have already been used to deliver multiple 10,000-accelerator clusters. The training of key versions of the Ernie 5.1 large model was completed on a fully domestic cluster, which also supports inference for more than 80 industry scenarios in the energy and power sectors, as well as training for in-house developed power industry large models.

Another group of vendors does not subscribe to the "bigger is better" philosophy.

The Alibaba Cloud Lingjun Zhenwu M890 uses a 64-accelerator unit as its core Scale-up building block, which can then be expanded into larger clusters via a Scale-out network. Unlike most super-nodes which are primarily designed for private deployment, Alibaba has also made this computing power form available on its public cloud, allowing customers to call on-demand resources without needing to purchase physical hardware.

Inspur Information and Muxi have also chosen a similar direction. Inspur Information launched a 64-accelerator super-node last year, followed by a 32-accelerator product this year, which can support inference for trillion-parameter large models. Zhao Shuai, Vice President of Inspur Information, explained, "We use trillion-parameter large models for internal AI Coding work, and the efficiency is absolutely better than using hundred-billion-parameter models. That's the value that trillion-parameter large models deliver. So we're launching this 32-accelerator product again to make super-nodes truly affordable for more enterprises." Muxi, meanwhile, bases its design on 64 accelerators per cabinet, with a focus on CUDA ecosystem compatibility and horizontal scalability to 10,000-accelerator clusters.

Some other vendors are attempting to break away from the NVIDIA GPU or Huawei accelerator paths, pursuing a third development route. WeCores uses a reconfigurable dataflow architecture, where each chip handles both computing and data forwarding, allowing chips to connect directly to one another without requiring any switches. Chen Yilun, Vice President of Products at WeCores, told Digital Intelligence Frontline that the company has already launched a 4096-accelerator super-node with a peak computing power of 500 PFLOPS. This system has been deployed at a training facility in Beijing, and has also seen commercial rollout in the financial industry and state-owned enterprises in Inner Mongolia, Xinjiang, and Zhejiang. East Computing Core, which released its products in July, is another vendor following this same technical path.

In addition, Moore Threads also plans to launch related super-node products alongside its new generation of GPUs by the end of the year. Cambricon does not currently offer super-node solutions, but industry insiders expect its next-generation chips will also support super-node architectures. At this point, nearly all major domestic computing chip and server enterprises have entered this market.

Why This Year?

Why has the super-node explosion happened precisely this year? Industry insiders point to three core driving factors.

First, demand for model training and inference has reached sufficient scale. On the training side, model parameter counts are skyrocketing. In July, Kimi released its K3 large model with 2.8 trillion parameters. "By next Spring Festival, model parameter counts will hit 3 trillion, then 6 trillion the following year, and eventually reach 10 trillion. Mid-sized overseas models already have 5 trillion parameters," Dr. Yang Jian, CTO of Muxi, told Digital Intelligence Frontline. The industry generally estimates that training a 1-trillion-parameter model requires roughly 5,000 accelerators, a 3-trillion-parameter model needs around 10,000 accelerators, and a 6-trillion-parameter model requires 20,000 accelerators for training. The optimal solution for this workload is to interconnect large numbers of chips into a single unified operating system.

On the inference side, a Huawei representative gave an example: in AI Coding scenarios, a single task can trigger tens of thousands of operation calls, while complex long-horizon tasks can involve hundreds of thousands or even millions of calls. "This translates to higher concurrency and longer context windows — exactly the core selling point of the super-node's large memory pool." Looking at the ratio of training to inference workloads, a senior Alibaba Cloud official told Digital Intelligence Frontline that the current ratio of online inference to training is close to 5:1 to 10:1, meaning "out of 10,000 accelerators, roughly 70% to 90% are running inference, with only a small fraction handling training."

Second, the economics now make clear business sense. While the price of a super-node is higher than the total cost of a set of regular servers with the same number of accelerators, customers are calculating total system efficiency. "You can't just look at the number of accelerators; this involves underlying infrastructure for power, cooling, and space, and super-nodes are far more resource-efficient. If all the big enterprises are deploying super-nodes, the math must definitely add up," a H3C representative told Digital Intelligence Frontline. Zhao Shuai, Vice President of Inspur Information, previously gave an example: "With super-nodes, if inference performance improves 10x while power consumption only increases 2x, you're still seeing net gains."

A Kunlunxin representative explained that in the traditional architecture, the combined computing power of 10 separate machines does not equal 10 — "it only adds up to 6". Super-nodes use high-speed interconnection to minimize this performance loss, ensuring more of the available computing power is put to productive use. For example, using super-nodes for PD-separated inference can deliver a 30% to 50% performance improvement over regular machines with the same number of accelerators.

Third, the industry has reached the stage where batch delivery is feasible. "Super-nodes are engineered products that integrate GPU chips, storage, networking, switching, thermal management, power supply, and software — no single vendor can develop all these components independently," the Kunlunxin representative said. "At this point, super-nodes have matured to the stage where they can be deployed in large batches." Another insider added, "Regardless of their specific performance levels, all vendors are now capable of delivering fully assembled units."

That said, are super-nodes better suited for training or inference? Different vendors hold different views on this question.

Huawei representatives believe super-nodes deliver value for both training and inference: trillion-parameter large model training requires massive memory pools, while Agent inference needs massive concurrent context capacity, meaning super-nodes must be equipped with high-capacity memory chips. Moore Threads argues that super-nodes should target training for 1-trillion to 5-trillion parameter large models, as well as high-speed inference "token factory" scenarios. Muxi representatives, meanwhile, state that "after rigorous verification, super-nodes deliver the most prominent value during the Decode phase of inference, when answers are generated token by token." Chen Yilun from WeCores analyzed that "currently, super-nodes are mostly used for fine-tuning on the training side, but they perform even better when deployed for inference workloads."

Divergences and Consensus

Amid this explosive growth, there are many disagreements over super-node technical routes, with no single standard answer yet — the final outcome will depend on each enterprise's own technical and business judgment.

The first major point of divergence is the scale debate: how large should a super-node be? 64 accelerators, 128 accelerators, or even bigger configurations? The industry is split into two camps, referring specifically to the Scale-up scope of super-nodes — that is, how many chips can be integrated into a single tightly-coupled computing domain using high-bandwidth, low-latency interconnection.

Some industry players belong to the "small-node camp". "In my personal view, 64 accelerators are enough. The maximum practical scale for TP parallelism and EP parallelism is around 64 accelerators, and the rest can be linked via PP and DP parallelism through Scale-out. There's no need to make the Scale-up of individual nodes infinitely large, because all communication introduces latency, and TP and EP are extremely sensitive to latency. Even if there's only a few microseconds of latency in a 1000-accelerator interconnection, the GPUs will end up idling and waiting," one insider told Digital Intelligence Frontline. "You have to account for both latency and stability. This is an engineering balance — bigger Scale-up isn't always better." Another industry figure shares the same stance, noting that "64 to 128 accelerators is the most cost-effective configuration, anything larger would be wasteful."

Other enterprises fall into the "large-node camp": Huawei's latest super-node scales to 1024 accelerators, Sugon's to 640, while Baidu Tianchi and H3C have both pushed their designs to 512 accelerators. "In the Agent era, with trillion-parameter models, million-Token context windows, and massive concurrent inference, the unified memory pool of 64/128 accelerator small nodes is simply too limited," they argue. "64 accelerators are more like a basic cabinet unit, suitable for small to medium model fine-tuning and offline inference. Full training of trillion-parameter large models and large-scale multi-agent concurrent workloads must rely on far larger super-node architectures."

The second major divergence is the CUDA compatibility debate. Vendors including Alibaba Pingtouge, Muxi, and Haiguang are pursuing CUDA-compatible routes, which lower the cost of code migration. Huawei Ascend, Kunlunxin, and other vendors, meanwhile, are taking the path of developing fully in-house independent ecosystems.

A major cloud enterprise insider analyzed to Digital Intelligence Frontline that while in-house developed ecosystems have higher initial migration costs, they offer greater controllability and more room for long-term evolution. He frankly stated that CUDA compatibility may no longer be a critical concern in two or three years. "Let's go back to first principles of technology. CUDA achieved its success because early vector programming imposed a huge mental burden on developers, and CUDA drastically reduced that burden, leading to its dominant position over the past decade. But now that Agentic coding is emerging, humans won't necessarily be the ones writing code — that work will be delegated to Agents, making CUDA far less essential