WAIC is packed with "shovel sellers", but computing power services are ultimately a game for the few.
At WAIC 2026, robots still grabbed the vast majority of the spotlight. However, the fastest-growing group in the exhibition hall consisted of companies branding themselves as "AI Infra" providers.
Many familiar faces populated the venue. Businesses focused on liquid cooling, switches, storage, and cluster software all added "AI Infrastructure" to their business cards this year. Resource integration platforms are also pivoting in this direction. Every link in the industrial chain seems to be moving closer to the user-facing entry point.
The logic is straightforward: with the monetization paths for large model applications still unclear, "selling shovels" to the AI industry appears to be a more certain business.
Yet stepping outside the exhibition hall, the real industrial landscape does not align with this bustling scene.
Over the past three years, numerous intelligent computing centers have been built across the country. Today, some are operating at full capacity with long queues of users, while others are publicly seeking clients at prices close to cost; meanwhile, model enterprises and research institutions continue to complain about insufficient computing power.
As more and more companies join the "shovel-selling" market, the amount of "usable computing power" available has not increased correspondingly.
The gap lies in a point that booth pitches never voluntarily mention: AI Infra is a complete industrial chain, where each company delivers only its own segment — chips, interconnection, storage, scheduling software. But what users ultimately purchase is not a single component, but a result — tasks completed on time and running stably. Organizing these scattered segments to deliver that result is what constitutes true computing power services.
The industrial chain can accommodate thousands of participants, but computing power services are destined to be a business for a select few.
To understand why, we must first recognize the unique nature of computing power as a commodity.
01 Advances in components do not automatically translate to usable computing power
The biggest misconception about computing power is that it appears to be a standardized product — billed by the card, settled by the hour, seemingly as ubiquitous as water and electricity.
This is far from the truth.
The same number of computing cards can perform drastically differently when deployed in varying network, storage, and software environments: a cluster optimized for large model inference may not withstand training workloads with high communication demands; a cluster that can run mainstream open-source models does not inherently support scientific computing or industrial simulation.
Even with identical hardware, variations in system scale, interconnection efficiency, task scheduling, and software adaptation will result in different amounts of effective computing power delivered to end users.
The industrial chain delivers component performance, while users require system-level outcomes. Between them lies a massive amount of engineering integration work — a segment that most participants in the chain do not engage in.
Consequently, a market thriving at the component level has spawned numerous awkward intermediate states: resources exist but are not user-friendly; platforms are available but lack control over underlying resources; clients are acquired but application problems cannot be resolved.
Matching platforms can inform users where idle cards are located, but cannot resolve driver incompatibilities, storage bottlenecks, or degraded cluster communication efficiency through a scheduling interface; resource owners can rent out equipment but cannot assist clients with application migration; system integrators can build the infrastructure but may not have the capability to continuously onboard tasks.
This explains why computing power shortages and resource idleness can coexist: what users lack is never just a single computing card, but a computing environment that is "ready-to-use, stable in operation, and with guaranteed fault recovery".
At this point, a natural question arises: don't large cloud vendors already provide exactly this kind of service?
In standardized scenarios, they certainly do. Leading cloud vendors' GPU cloud services are mature, with comprehensive elasticity, billing systems, and ecosystems.
However, cloud business models are built on standardization and scalability, inherently prioritizing customers with large, predictable workloads and clear profit margins. Several types of demands fall outside the coverage of this model:
Scientific computing and industrial simulation feature high customization, limited scale per customer, and requirements for FP64 precision and specialized software stacks, resulting in far lower return on investment compared to standard inference services. Government, enterprise, and research clients with data sovereignty, local deployment, and information technology application innovation requirements do not just need a tenant account on the public cloud, but a system built in their own data centers with long-term dedicated support. Heterogeneous resource integration across different chips and data centers even directly conflicts with cloud vendors' business logic of "keeping customers within their own technical ecosystem".
The demands that clouds struggle to meet — comprehensive, heterogeneous, long-cycle, service-intensive — are the hard nuts that computing power services must crack. The challenge is not to put computing power online for sale, but to organize scattered computing resources into sustainably deliverable capabilities.
02 Behind a single computing card, there are three sets of accounts to balance
The viability of this business hinges not on graphics card prices, but on the ability to balance three sets of accounts simultaneously: construction accounts, operation accounts, and customer accounts.
First, the construction accounts —
Users expect to purchase computing power on demand like water or electricity, but service providers face long-cycle, asset-heavy projects: data centers, servers, networks, storage, liquid cooling, and power supplies all require upfront investment, with major equipment depreciated over 4 to 5 years. Yet clients may only rent resources for weeks or even days.
Mismatched timelines are the fundamental reality of this business.
Worse still, the update cycle for AI chips and system architectures has shrunk to about one year. Before old equipment can recoup its costs, new-generation products have already hit the market — technological iteration outpacing depreciation is the primary risk in computing power operations.
Asset-light platforms seem to avoid this issue: they focus on leasing and matching without holding hardware inventory. But the risk does not disappear; it is merely transferred to equipment owners. When supply is tight, platforms cannot guarantee resource availability; when the market cools, they do not share the idle costs with upstream partners.
Next, the operation accounts —
As scale expands, failures are no longer exceptions but daily occurrences.
A public reference point: Meta disclosed during the training of Llama 3 that a 16,000-card cluster experienced over 400 unexpected interruptions across a 54-day training period — an average of once every 3 hours, mostly caused by GPU and memory hardware failures.
This is the performance delivered by one of the world's most capable engineering teams in a homogeneous environment. Scaling up to 100,000-card clusters with heterogeneous chips and mixed workloads of training, inference, and scientific research will only introduce more risk variables.
Therefore, genuine operations go far beyond simple computing power allocation. It requires continuous management of resource orchestration, task prioritization, fault isolation, dynamic migration, and system recovery, while preventing certain tasks from monopolizing resources and slowing down all other users. These tasks demand deep access to the underlying system. Companies that only control a platform entry point without access to networks, storage, and computing environments cannot provide reliable guarantees for "when tasks will finish running".
Finally, the customer accounts —
On the day a computing center is completed, equipment does not automatically generate revenue. The real test begins after delivery: can enough well-structured tasks be secured to maintain utilization at a sustainable level?
The challenge lies in the vast diversity of task types. Large model training consumes resources extremely rapidly, and demand plummets once a project concludes. Inference workloads are relatively continuous but are highly sensitive to response speed and cost. Scientific research tasks run for months with strict requirements for precision, network, and storage. Industrial clients prioritize data security, local deployment, and industry software compatibility.
A long-running computing platform must integrate these diverse demands into a comprehensive scheduling table: prioritizing critical tasks during peak hours, onboarding high-throughput jobs during off-peak periods, and leveraging workload complementarity to smooth out usage fluctuations.
Sales entry points can be lightweight, but utilization can only be gradually optimized through robust customer systems, application migration capabilities, and model adaptation expertise.
03 Only when three sets of accounts are managed under one unified system can it be called computing power services
Using these three accounts as criteria to re-examine the AI Infra market, a clear dividing line emerges: most participants only focus on one of the accounts, while very few entities can holistically manage all three within a single system.
What do the latter look like? They possess three highly replicable characteristics.
First, they have a large-scale on-site engineering team. Project completion is not the end of delivery; engineers are permanently stationed at client sites and on the front lines of systems to handle network jitter, equipment failures, software upgrades, and application migration. These teams are recorded as costs on financial statements, but represent service commitments to clients.
Second, they dare to make long-cycle infrastructure investments. They can navigate construction, ramp-up, and technology transition phases, rather than evaluating short-term quarterly returns for individual product lines.
Third, they control the critical links that determine service quality. They do not need to dominate the entire industrial chain, but core capabilities that determine "whether tasks can complete" — including networking, storage, scheduling, and software adaptation — must be firmly in their control or within their efficient coordination scope.
Measured against these standards, qualified entities in China are few and far between: a handful of computing power enterprises with systems engineering capabilities, and telecom operators that own nationwide networks, data centers, and government-public service systems. These two types of capabilities are complementary — the former excel at building, optimizing, and deploying complex computing platforms, while the latter boast nationwide infrastructure and billing/customer service organizations.
What truly makes them rare is their lack of escape routes: if utilization is insufficient, they bear the costs themselves; if systems fail, their teams step in; if client tasks cannot run, they cannot shift responsibility to downstream suppliers.
This is the watershed between computing power services and the computing power supply chain.
04 The industrial chain can be divided, but responsibility cannot be fragmented
The concentration of computing power services among a few entities does not mean other players are pushed out. On the contrary — the more complex the system, the higher the value of each specialized segment in the chain.
As cluster scales expand, scheduling, monitoring, metering & billing, automated operations, and performance optimization become more valuable; as heterogeneous resources increase, professional platforms are needed to shield users from underlying technical complexities; as industry clients flood in, service providers with deep domain expertise are required to handle model deployment and application migration. When these segments are deeply developed, they become irreplaceable positions.
The only real question remains: who will orchestrate this entire chain?
The foreseeable landscape is that system-level entities bearing ultimate responsibility will act as "chain leaders", taking overall charge — ensuring system stability, guaranteeing task delivery, and delivering exceptional customer service; software platforms, data centers, system integrators, and industry service providers will excel in their respective segments, integrating into the overall delivery system through standardized interfaces and clear responsibility agreements. Responsibility is clearly assigned, and a collaborative ecosystem of division of labor is formed.
For most companies, becoming an indispensable link in this chain is far more valuable and less risky than building their own resource trading platform from scratch.
Policy orientation aligns with this direction. National integrated computing network documents have explicitly proposed fostering professional computing and network operation entities, and improving systems for resource scheduling, demand matching, metering, billing, trading, and settlement. Notably, the keyword is "professional operation entities", not more resource entry points.
This reflects a shift in evaluation criteria: in the construction phase, the industry competes on equipment count, peak performance, and cluster scale; entering the operation phase, utilization rate, task completion rate, fault recovery time, application coverage, and unit computing cost will become the new scorecards.
Whoever can consistently deliver stable and usable computing power will truly win this market.
WAIC 2026 brought unprecedented hype to AI Infra, but this excitement is mostly concentrated on the supply chain side. As the industry begins to carefully calculate depreciation, utilization, and delivery efficiency, many companies will eventually return to their most proficient segments. The few enterprises that can both build robust systems and willingly assume long-term operational responsibilities — along with the clearly divided service ecosystem that grows around them — will ultimately remain at the center of the stage.
This article is from the WeChat public account "Bohu Finance", authorized for release by 36Kr.