HomeArticle

Has the biggest shortcoming of domestic computing power been filled by a Token factory?

晓曦2026-07-20 19:34
SenseNova Supercomputing Facility attempts to integrate chips of different brands and generations into a continuously profitable Token factory.

In the past two years, the most frequently discussed unit in AI infrastructure circles is the P.

For example, how many P of computing power a company owns, or how many GPUs a smart computing center is equipped with, has instantly become the most intuitive way to judge its strength. Local governments building smart computing centers and large model companies purchasing computing power mostly focus on the number of chips and the scale of computing power.

However, after entering the Agent era, this calculation method is changing. According to IDC data, in January 2025, the daily average Token consumption in China's enterprise-level MaaS market was about 1.6 trillion, which skyrocketed to 9.6 trillion by the end of December of that year; the total Token consumption in 2026 is expected to reach 40,000 trillion, about 20 times that of 2025, equivalent to a daily average of nearly 110 trillion.

The underlying growth momentum mainly comes from the continuous maturation of multimodal models and the large-scale implementation of Agent applications. The functions and roles played by Agents are gradually enriching, far beyond passively answering questions. They need to read longer contexts, invoke search, databases and external tools, break down tasks and conduct multiple rounds of reasoning.

In response, Zhang Xingcheng, Chief Scientist of SenseCore Business Group, said in an interview that in the past, models for chat scenarios had short inputs and long outputs; in Agent scenarios, the input-output ratio could reach 100:1 or even higher, and the context has also jumped from tens of thousands of Tokens to hundreds of thousands or even millions of Tokens.

Therefore, what enterprises are actually purchasing is no longer abstract computing power, but the Tokens that computing power can ultimately produce. This has led enterprises to start thinking: how many Tokens can be generated with the same budget? How much inference can be supported by the same kilowatt-hour of electricity? How long does it take to migrate a model to a new chip? Can it remain stable under high concurrency? These are inevitably becoming the new ledger of computing power.

36kr learned that during WAIC 2026, SenseCore respectively launched heterogeneous hybrid inference, full-stack adaptation, computing-power-electricity coordination Agent and domestic ecosystem plan. It attempts to answer the question: as domestic AI chips become more and more numerous, how to organize chips with different architectures and different performances into a stable, low-cost, and even profitable Token production system? All these actions point to the positioning of "the AI infrastructure that best understands large models".

Domestic computing power crosses the profit line not by relying on a more powerful single card

In the past, the primary problem faced by domestic computing power was "whether it exists". As domestic AI chip companies have successively entered the commercialization stage, the problem is transforming into "how to use it". Different chips adopt different hardware architectures, compilers, operators, communication methods and software toolchains. A model that has been running on the NVIDIA platform often needs to be re-adapted when migrated to a domestic chip. Switching to another chip means repeating all the pitfalls again.

Zhang Xingcheng calls this work a "toilsome task". SenseCore started domestic chip adaptation in 2018, gradually integrating chip original manufacturers, SenseCore R&D teams, universities and scientific research institutions into the same system, and promoted the DeepLink open source system with Shanghai AI Laboratory to uniformly adapt to operators, compilation and heterogeneous communication of different chips.

At the same time, Agents are also being used to solve the problems of the infrastructure itself. The operator migration, development and performance tuning that used to be completed one by one by engineers can now be assisted by AI. After chip manufacturers provide underlying operators, SenseCore performs automated migration and optimization on the platform, trying to reduce the engineering cost of domestic chips entering new models and new businesses.

But adaptation is only the first step. In the first half of this year, SenseCore began to push a previously verified technology to large-scale commercial services: heterogeneous hybrid inference. Large model inference can usually be divided into two stages: Prefill and Decode. Prefill is responsible for reading and understanding a large amount of context input by users, which is computing-intensive; Decode generates Tokens one by one, which has higher requirements for video memory bandwidth, latency and stability.

In traditional solutions, these two stages are often completed by the same type of GPU. The idea of SenseCore is to split the tasks: let general-purpose domestic chips undertake Prefill, reserve high-end chips for Decode, and connect the two through a unified network, scheduling and cache system. Through PD separation, heterogeneous networking and system optimization, higher overall Token output can be obtained at lower hardware costs.

In addition, unlike other Token factories, the realization of such combinatorial optimization and technical implementation also relies on the super-large-scale AIDC smart computing center self-built and owned by SenseCore. Xuan Shanming, CTO of SenseCore Business Group of SenseTime, said that the optimization of domestic computing power is not just about migrating models to a domestic card. Relying on the self-built Lingang AIDC, SenseCore can simultaneously transform the hardware combination, network architecture and software system, and try different heterogeneous combinations of domestic chips and high-end chips. "Each individual part may not have an exceptionally outstanding advantage, but when combined, they form a platform-level long board." In contrast, many computing power operators do not have the right to physically transform the computer room, and software manufacturers cannot easily connect hardware and network systems, making heterogeneous hybrid inference difficult to truly implement.

The key to this set of combinatorial optimization lies precisely in not pursuing the best performance of a single card in all tasks. Yang Fan, Co-founder of SenseTime and President of SenseCore Business Group, used the triathlon as a metaphor: a swimmer may not achieve good results in a triathlon alone, but if three athletes who are good at different events are allowed to take over in relay, the overall result may be higher.

The same logic applies to AI infrastructure. Large language models, video generation, world models, embodied intelligence and AI for Science have different requirements for computing, video memory and communication. Even if higher-end domestic chips undertake Decode in the future, older-generation chips can still be responsible for Prefill or data generation, and new-generation chips undertake more complex tasks. Heterogeneous combinations will not disappear as a result.

In the view of SenseCore, heterogeneity is not a stopgap measure when domestic chips are not yet mature, but a way to find the system optimal solution when models and hardware are iterating rapidly and demands have not yet converged. Its commercial results have begun to show. According to Yang Fan, domestic hybrid inference clusters can currently achieve a positive gross profit margin.

The Token service volume of SenseCore is also expanding. According to its estimates, the daily average service volume will rise from about 400 billion Tokens at the beginning of 2026 to about 2.42 trillion at the end of July, with the year-end target reaching 10 trillion, about 25 times that of the beginning of the year.

Xuan Shanming also admitted that the commercialization of domestic computing power will not occur simultaneously in the training and inference links. At present, domestic chips still face great challenges in model training, and relatively mature closed loops mostly appear in models such as video, image and text. In contrast, inference is more likely to expand first: through PD separation and heterogeneous hybrid inference, different chips are placed in tasks they are better at, which can not only improve the overall throughput, but also drive domestic chips into real businesses.

SenseCore's positioning of "best understanding of large models" is also built on this task decomposition capability. In Yang Fan's view, technical knowledge spreads quickly, and a single capability is difficult to maintain for more than two years. A more lasting barrier is whether the organization can always be one step faster than others, understand where downstream models will evolve, and then adjust the infrastructure in reverse. Compared with other AI enterprises, SenseCore has a unique organizational advantage: the model team and the infrastructure team are close enough. For example, Zhang Xingcheng, Chief Scientist of SenseCore, also participates in the work of SenseTime Research. Through job rotation and joint R&D, the two teams quickly transfer changes in model load to chip selection, network, cache and scheduling systems.

This end-to-end capability also extends to the power system. SenseTime Lingang AIDC was rated as the first 5A-level smart computing center in China in 2024. This certification comprehensively examines indicators such as theoretical computing power, effective computing power, computing power energy efficiency and business scenario support capabilities.

Traditional data centers usually use PUE to measure energy efficiency, but PUE can only reflect how much electricity is used for IT equipment, and cannot answer how much effective Tokens are ultimately produced by that electricity. In extreme cases, to optimize PUE, the computer room may make the server fans run at high speed, but the GPU reduces the frequency due to temperature or power consumption restrictions. The local indicators are better, but the actual computing power output decreases instead. SenseCore therefore proposed TPW, that is, the Token output corresponding to the unit power cost.

In Lingang, SenseCore connects the energy storage system, computer room management system and computing power operation platform, and schedules computing power according to task real-time requirements, electricity price and grid load. Offline inference, model evaluation and some pre-training tasks do not need to be executed immediately, and can be shifted to time periods with lower power loads. For example, when the Shanghai power grid carried out demand response in July, the Lingang data center released 75% of its power consumption capacity within two hours, reducing power consumption by about 46,000 kWh. SenseCore hopes to prove through this that smart computing centers are not necessarily just large power consumers, but can also become a flexible "capacity pool" in the power grid.

However, electricity bills are not yet the decisive factor for the profitability of China's Token business. Yang Fan judged that computing-power-electricity coordination can create a cost difference of about 10 percentage points in China, and its value may become more prominent in overseas markets such as Southeast Asia and the Middle East where electricity prices are higher.

Three-pronged approach covering regions, scenarios and ecosystems, making Token factories accessible to everyone

In fact, there is no shortage of companies building smart computing centers in China, but the participants of a large number of projects are layered and fragmented. Some provide computer rooms, some purchase servers, some are responsible for networking, some underwrite computing power, and some do retail for end customers. Each layer needs to make a profit, but no single entity can see all the data, let alone conduct joint optimization among models, chips, networks, power and customer tasks.

Yang Fan believes that only by truly accessing end-user traffic and scheduling according to different customers, models and tasks, can it be possible to find the optimization space for Token costs.

This is where the difference between SenseCore and large manufacturers emerges. Large manufacturers such as Alibaba and Volcano Engine also have end-to-end capabilities from data centers to models and applications, but their highest-quality resources usually prioritize serving their own e-commerce, content and cloud businesses. SenseCore, on the other hand, is willing to provide high-stability, high-elasticity, scalable end-to-end system services for AI companies, scientific research institutions and industrial customers.

According to the target disclosed by Yang Fan, SenseCore plans to build at least 5 domestic clusters with ten thousand cards, serving more than 200 AI startup organizations. Choosing 5 ten-thousand-card clusters with different combinations instead of building a single-architecture one-hundred-thousand-card cluster is precisely to retain flexibility for different models and scenarios.

Regional projects undertake the task of verifying business models. In Yancheng, the first phase of the SenseTime Smart Computing Center is planned to have 3000P of computing power, and it tries to form a closed loop from infrastructure to local advantageous industrial applications by introducing AI enterprises, industrial chain resources and funds. The Hong Kong, China project plans to complete the first batch of computing power clusters by the end of 2026, and build a smart computing center with a computing power scale of 40,000P by 2030.

The Saudi project is still in an earlier stage of market verification. Yang Fan also admitted that the focus of AI infrastructure going overseas in 2026 is still on layout and path validation. There are only a small number of projects in Southeast Asia and the Middle East at present, and larger-scale revenue may come in the next stage.

In addition to launching projects in Saudi Arabia and Hong Kong, China, SenseCore is also accelerating the exploration of a large number of more innovative and imaginative scenarios, so as to continuously expand the boundary of domestic computing power, such as space computing power. At the 2026 WAIC conference, SenseCore and SpaceStar Aerospace jointly built the SenseTime Computing Power Constellation, launching satellites for preliminary verification. Although in the next two years, space computing power will be more in the stage of technical reserve and pre-research, it still has high potential value in China, such as solving AI service problems in overseas weak network areas and satellite coverage scenarios.

In order to expand technical capabilities to more domestic software and hardware, SenseCore also joined forces with chips, core components and AI Infra to launch the "Galaxy Plan". This plan aims to solve two problems: first, to make domestic production not only stay at the GPU level. Even if a server uses a domestic GPU, the CPU, memory, optical modules and switches may still come from overseas, and there are also adaptation and joint debugging problems between domestic CPUs and GPUs. The hardware supply chain needs to be jointly verified by software, models and real applications, otherwise the so-called "full domestic production" can easily stay on the equipment list.

Second, to find continuous demand for domestic computing power. Companies engaged in embodied intelligence, world models, video generation and AI for Science need computing power, but may not have the ability to complete chip adaptation, cluster operation and maintenance and performance optimization on their own. SenseCore hopes to provide them with computing power, Tokens and tools, and then bring these companies into urban, scientific research and industrial intelligent projects.

Dongwu Securities proposed that 2026 is the "first year of large-scale commercial use of domestic AI infrastructure". From the demand side, the Token market is exploding, and the gap between supply and demand is expanding; DeepSeek V4 actively and deeply adapts to domestic computing power, laying a foundation for the accelerated rise of domestic computing power.

And a domestic AI infrastructure service provider, with heterogeneous hybrid inference and end-to-end service capabilities, has made more people see the dawn of domestic computing power.

However, for the entire domestic AI infrastructure ecosystem, there is still a long way to go in the future. The bigger bottleneck lies in how to make domestic clusters truly run stably for a long time, continuously reduce Token costs, maintain a sufficiently high utilization rate of smart computing centers, and ensure that customers are willing to renew their subscriptions continuously.

For SenseCore, they hope that "the AI infrastructure that best understands large models" is not just a positioning, but can be translated into a clear economic ledger: every card, every cabinet and every kilowatt-hour of electricity can ultimately be transformed into high-quality Tokens that many people are willing to pay for.