Verified by trillion-level Token call volume, how does PPIO build a new generation of intelligent Token factory?
The Token Economy Is in Full Swing
The AI industry in 2026 is undergoing a handover of dominance from training to inference. At a press conference held by the State Council Information Office in March, China's National Data Administration released a set of figures: as of March 2026, China's average daily Token call volume has exceeded 140 trillion, representing a more than 1,000-fold increase compared with 100 billion at the beginning of 2024. In just two years, this figure has grown by four orders of magnitude.
Token is no longer a rare unit of measurement in technical documents. It has become the most basic transaction unit in the AI era — just like electricity, water, and the CPU core count in the past, with the difference that its expansion rate far exceeds that of any other basic resource in human history.
IDC also predicts that the global annual Token consumption will jump from 0.0005 Peta Token in 2025 to 150,000 Peta Token in 2030, with a compound annual growth rate as high as 3418%. By 2031, the number of global active agents will reach 350 million.
What drives this outbreak is exactly the structural shift of the AI track.
Before 2026, the main battlefield of the AI industry was training, where enterprises stacked GPUs, ran parameters, and competed for rankings. But after 2026, inference is becoming the main engine of computing power consumption. IDC estimates that the scale of China's AI server market will reach 3.5 trillion yuan in 2026, the demand structure has shifted from "training-driven" to "training + inference dual-wheel driven", and the shipment volume of inference servers is close to that of training models.
In addition to the continuous expansion of inference demand, the emergence of agent applications and the cliff-like drop in inference costs are all pushing the industry to an outbreak. More and more consensus is emerging in the industry: AI is evolving from "answering questions" to "getting things done". At this year's WAIC, multiple agent applications were launched alongside new mobile phones, and every autonomous planning, tool call, and multi-step execution of agents consumes Token in an exponential way.
The impact of this transformation on computing power infrastructure is also subversive.
Yao Xin, co-founder and CEO of PPIO, told 36Kr, "Traditional cloud computing serves human beings. Programmers will use a virtual machine for at least several days or even months; while Agent call tasks are fragmented and high-frequency. The minimum billing unit for Sandbox on the PPIO platform is already counted in seconds".
When people use the cloud, there are peaks and valleys that follow the cycle of work and sleep. But Agents use the cloud 7×24 hours.
In terms of latency, there are also differences between the two scenarios. Humans have a latency tolerance at the second level, but when an Agent completes a complex task, it needs to call dozens or even hundreds of times repeatedly. After the latency of hundreds of milliseconds per call is amplified by the loop, the task execution efficiency will become unacceptable.
These differences mean that the cloud computing architecture designed for humans cannot directly serve Agents.
Against this background, the Token factory has become an indispensable link in adjusting computing power capacity and has also attracted much market attention.
According to data from CIC, based on the average daily Token consumption in 2025 and the first quarter of 2026, PPIO ranks first among independent AI cloud computing service providers in China. In April 2026, the average daily Token consumption of the platform reached 1.03 trillion; by June, this figure further exceeded 1.2 trillion, an increase of more than 8 times over the same period in 2025.
According to another disclosure, PPIO's AI cloud computing revenue jumped from 10.387 million yuan in 2024 to 119.2 million yuan in 2025, a year-on-year increase of more than 10 times. The number of global registered developers on the platform increased from 125,000 at the end of 2024 to more than 666,000 in June 2026.
Solving the Efficiency Problem, the Smart Token Factory Is Here
When the inference cost drops 10 times every year, the business model of simply providing computing power resources is rapidly losing its premium capability. What is truly valuable is who can deliver Token with higher efficiency, lower cost and better experience.
The "Smart Token Factory" defined by PPIO is a large-scale production and delivery system optimized around the full life cycle of Token.
Yao Xin told 36Kr that although the concept of Token factory emerged at the GTC conference in March 2026, PPIO has been working on inference services since 2023 and launched the MaaS platform in 2024, accumulating more than three years of experience in this direction.
"Today's Token factory is nothing rare. What matters more is how to continuously improve the intelligence level of the Token factory." Yao Xin believes, "You buy a bunch of computing power and deploy models, which seems to be able to produce Token. But the problem is, how to produce smarter Token with high quality and high efficiency?"
Therefore, Yao Xin put forward a core formula for the Agent era — Agent Productivity = Token Intelligence Density × Agent Loop Duration. Among them, the Token intelligence density determines the upper limit of the quality of each step of Agent decision-making, and the Agent Loop duration determines how long the Agent can run continuously and how complex the tasks it can complete.
Around this formula, PPIO Smart Token Factory has taken the lead in launching a smart model gateway in China, which is the intelligent scheduling center for AI Agents — key decisions are checked by hybrid models to improve quality, and simple tasks are automatically diverted to lightweight models through model scheduling, ensuring that the Agent completes tasks with the lowest Token cost and the highest intelligent performance, and continuously increases the Token intelligence density.
Regarding the hybrid model, Yao Xin revealed to 36Kr that PPIO is testing the combined use of two to three different models, letting them compete and discuss with each other. In internal tests, this hybrid model method has been able to outperform GPT-5.6 in some task executions. "Although there may still be a small gap in the performance of a single model, by paying a certain amount of computing power and Token consumption, we can obtain stronger final task execution capability." This means that the intelligent upper limit of the model may not only lie inside the parameters, but can also be broken through from the outside through engineering methods.
This is the working principle of the smart model gateway. By providing a hybrid inference system for AI applications, the smart model gateway turns each model call into a "expert consultation".
In the past, PPIO has built a full-stack AI cloud capability from the bottom up, covering GPU clusters, inference service optimization, and in-depth understanding of application scenarios, which reflects the ultimate efficiency capability.
Yao Xin gave us examples of three typical scenarios: human dialogue communication, agent calls, and AI programming, which have completely different requirements for performance parameters. Humans pursue second-level response, agents require millisecond-level response, and AI programming has the highest requirement for accuracy, hoping to make fewer mistakes. Customized optimization based on different scenarios is an important difference between PPIO and other platform companies.
In terms of specific product capabilities, PPIO's differentiated route is also reflected in several dimensions:
Quick integration, out-of-the-box availability.
Platform neutrality, not bound by any model ecosystem. Yao Xin revealed to 36Kr that an average application developer needs to call 10 to 15 different types of models, and has to switch according to model iterations every three months. The PPIO platform supports the unified access of more than 200 open source models, and users only need to modify one line of code to switch models. This neutral positioning has built trust among developers and is also an important strategic choice to enter the market.
Self-developed inference acceleration engine, deeply optimized for the Agent call mode. Different from the general architecture of traditional cloud computing, the entire technical stack of PPIO adapts to the fragmented, high-frequency and continuous characteristics of Agent loads from the bottom layer.
It is worth mentioning that the average GPU utilization rate in the industry is usually between 40% and 50%, while PPIO has long maintained it above 75%. Yao Xin told 36Kr that relying on distributed computing power scheduling capabilities, the company has connected more than 5,000 computing power nodes across six continents around the world, and relies on time stagger scheduling between the eastern and western hemispheres to integrate loads of different time, space and request frequencies with high quality, finally presenting an almost straight utilization curve in the background.
Nowadays, the cost of computing power is rising continuously. Whether the already produced computing power can be "fully utilized" directly determines the profitability of enterprises in the future. PPIO's performance on this indicator constitutes its important moat.
Agentic Cloud: The Next Infrastructure Wave of Token Consumption
The Smart Token Factory solves the problem of "how to efficiently produce Token today". And Agentic Cloud answers the more grand proposition of "who will consume Token tomorrow and in what way".
In 2026, global cloud giants have simultaneously released their Agentic Cloud strategies. Google put forward the Agentic Cloud strategy at the Next'26 conference, and launched Agent Engine and Agentic Data Cloud; Alibaba Cloud announced the completion of full-stack agent-oriented upgrade; Amazon launched AgentCore; Microsoft's Foundry Agent Service was officially commercialized. There is only one goal, which is to make Agents the core audience of the cloud.
At WAIC 2026, PPIO officially released the new positioning of "Agentic Cloud". Yao Xin believes that the core logic of this strategy is: The primary user of the cloud is shifting from humans to AI Agents.
The industry trend is also undergoing such a transformation. The autonomous inference, tool call and multi-step workflow of AI agents have the characteristics of continuous, high-frequency and intensive Token consumption, which puts forward new requirements for cloud infrastructure.
Yao Xin also shared a set of data with 36Kr: among the 8 billion people in the world, maybe only 20% have used AI, and only about 5% have paid for it, which means that there is still huge room for growth in human use of AI. But at the same time, the call of agents is growing exponentially. Taking PPIO as an example, there are more than 300 employees, and there are nearly 1,000 Agents in the background, and work such as development, operation and maintenance, and customer service are all being agentized. One person can own 10 or 100 Agents, and each Agent consumes Token 7×24 hours.
Based on this diverse demand, PPIO divides the product architecture of Agentic Cloud into three layers: infrastructure layer, model service layer and Agent Harness platform layer.
It needs to be specially noted that Harness, as an engineering framework specially used to constrain, guide, verify and correct Agent execution, covers all links except the large model itself: context construction, tool orchestration, verification loop, cost control and observability.
Sandbox is one of the core components of Harness. Last year, PPIO launched the first domestic Agent sandbox compatible with the E2B interface, providing a safe and isolated operating environment for Agent Harness.
Yao Xin mentioned that when the sandbox was released last year, the industry's understanding of agents was still limited to "the hands and brains of model calls". At the beginning of this year, with the popularization of open source agents such as OpenClaw, people realized that the real pain point is the continuous execution of long-term tasks and complex tasks.
According to official disclosure from PPIO, the cold start latency of PPIO Agent sandbox is less than 200 ms, with system-level security isolation, allowing each Agent task to run in an independent virtual machine environment; it supports the simultaneous creation of tens of thousands of sandboxes, and billing can be automatically suspended when tasks are idle, the comprehensive usage cost is more than 90% lower than similar products. Since the launch of PPIO Agent sandbox for one year, its business scale has increased by more than 120 times.
Nowadays, the order of Token pricing has produced its own rules: programming is the most expensive, followed by agents, and the cheapest is the dialogue service for humans. This order itself indicates the direction of future Token consumption.
Yao Xin believes, "The objects we serve have changed from humans to Agents, and even robots in the future."
This transformation is reshaping the business model of cloud computing. In the past, cloud computing competed on the number of functions and the degree of ecological closure. But in the Agent era, the competition lies in who can make silicon-based life run faster, cheaper and more stable. PPIO's choice is: not to build a closed garden, but to embrace open source; not to replace the full stack, but to be a complementer in the AI era.
The Token economy is redefining the value scale of computing power, and in this reconstruction, only those with the highest efficiency can get the final ticket.