Sino-US Token Economics: Profit Sources, Premium Flow and Realization Sequence
Profits in the large model industrial chain are being realized sequentially in the "upstream, midstream, downstream" order. The upstream computing power base has taken full advantage of the scarcity dividend first, while the midstream model layer is mired in price deflation as open-source capabilities catch up.
In terms of premium flow, China and the United States show significant divergence: the incremental AI value in the US is deposited in the existing high-priced software subscription system; in China, due to differences in payment willingness, the low-price Token dividend directly overflows to the downstream application layer.
In the future, whether downstream manufacturers can retain profits will depend on the "migration cost" they build in business scenarios.
The mismatch between the skyrocketing computing power of large models and the slow monetization is reshaping the profit distribution pattern of the global AI industrial chain.
In the past two years, the daily average Token call volume in the Chinese market has surged more than 1,000 times, but the total annual revenue of public cloud MaaS (Model as a Service) in 2025 only stayed at the level of 3 billion yuan. Massive consumption has not been converted into equivalent book income, and China and the United States have taken completely divergent paths in computing power bottlenecks and commercialization paths.
Song Xinzhu, an analyst at Northeast Securities, pointed out in the analysis of the Token economy industrial chain that AI profit accumulation consists of four mechanisms: scarcity premium, generational premium, integrated internal settlement income, and migration cost premium.
At present, profits are entering the financial statements in the order of upstream, midstream and downstream:
The upstream computing power base takes the lead in realizing the scarcity dividend;
The midstream model layer is mired in deflation caused by the commoditization of capabilities of the same generation;
The downstream application layer receives the dividend of falling computing power prices, and builds a long-term moat with the "migration cost" accumulated over time.
In the final outcome of the premium flow, constrained by the differences in market payment endowments between the two countries, the incremental AI value in the United States is being deposited in the high-priced software subscription system, while the low-price Token dividend in the Chinese market directly overflows to the application layer, waiting for the value revaluation after the full migration of the pricing method.
01
Computing power investment is approaching the cash flow boundary, and the thousand-fold call volume only supports a 3 billion-yuan market
The Token economy is still in the heavy asset construction phase
On the demand side, China's daily average call volume has soared from about 100 billion times in early 2024 to 100 trillion times at the end of 2025.
However, most Token consumption occurs within the in-house scenarios of large manufacturers, and no external transactions are formed; for the part of external transactions, the transaction price has been extremely compressed; in addition, the charging of the application layer has not fully migrated to Token pricing, resulting in a thousand-fold increase in usage only bringing a public cloud MaaS market size of 30.7 billion yuan.
Corresponding to the meager API revenue is the extremely heavy load on the computing power investment side.
The capital expenditure intensity has approached the coverage boundary of operating cash flow. As of the second quarter of 2026, the ratio of TTM (Trailing Twelve Months) capital expenditure to operating cash flow of the four major U.S.-listed cloud manufacturers rose to 0.63 to 1.05. Alphabet recorded negative free cash flow for a single quarter for the first time, and Meta's free cash flow plummeted 91% year on year. The source of funds in the construction phase has expanded from operating cash flow to the capital market.
The investment rhythm of the Chinese market is clearly differentiated: Alibaba's capital expenditure intensity ranks first among Chinese concept stocks, while Baidu is the only company among the eight leading computing power buyers that sees both revenue decline and increased capital expenditure.
02
The upstream takes full advantage of the scarcity dividend, and Sino-U.S. computing power bottlenecks diverge
The upstream is the only link that currently stably records profits in financial statements. The "scarcity premium" based on supply gaps directly generated NVIDIA's data center revenue of 193.7 billion U.S. dollars in FY2026. Faced with the same hunger for computing power, China and the United States have formed completely different clearance methods and industrial bottlenecks under the same regulations.
The industrial chain bottleneck in the United States lies in power access. In the queue of ERCOT (Electric Reliability Council of Texas) waiting for access approval, more than 90% of the over 1,800 projects are data centers, corresponding to a cumulative power demand of about 474GW. The extended cycle of approval and power access has pushed the North American data center vacancy rate to a historically low level. Scarcity in the United States is finally cleared by price, and the price increase revenue goes to leading manufacturers such as NVIDIA.
The industrial chain bottleneck in China points directly to computing power chips. Under export control, the Chinese market is cleared in accordance with regulatory allocation, which institutionally drives domestic substitution. In 2025, local manufacturers have accounted for more than 40% of the AI accelerator card market. Newly added computing power is concentrated in the "East Data, West Computing" hub nodes, and the construction entities include public sectors, operators and private capital, forming a public sector-led computing power system.
03
Open weights break through generational barriers, and midstream models become standardized production capacity
Tokens of the same capability level have extremely strong substitutability, and open weights (open source) have become the absolute main force to flatten the price gap. The cost for purchasers to switch suppliers is extremely low, and competition directly focuses on the listed price. Calculations show that the call price to reach the same capability as GPT-4 drops to about 1/40 of the original level every year.
Prices of capabilities of the same generation are rapidly converging globally. At the level of about 51 points on the AA Intelligence Index, the mixed prices of the four leading models in China and the United States (GPT-5.6 Luna, GLM-5.2, MuseSpark 1.1, Gemini 3.6 Flash) all fall into the very narrow range of 14 to 22 yuan per million Tokens. The lowest price at this level does not come from Chinese manufacturers, but from Meta, which enters the market in the form of API. Once the capability level is caught up by open source, Tokens become commoditized, and prices only change with usage volume and costs.
The midstream is thus squeezed from both ends. On the selling price side, the actual transaction price is often about an order of magnitude lower than the listed price; on the cost side, more than 70% of the unit inference cost is depreciation and amortization, and the electricity bill accounts for less than 10%. The core space for cost reduction does not lie in electricity prices, but in the depreciation period and computing power utilization rate.
In China, Token commoditization is being promoted to the infrastructure level. By directly subsidizing the purchasing side through computing power vouchers and establishing a unified measurement and price comparison platform, the intermediate markup space is squeezed out. The circulation sector is destined to see increased volume but thin profits. Only two ends can make sustainable profits: the production end earns the cost difference by extreme utilization rate, and the scenario end retains profits by customer switching cost.
04
Divergence of profit destination: value in the U.S. is deposited in subscriptions, while in China it overflows to the application layer
The deflating Token price releases the dividend of model upgrading to the downstream, and the way China and the United States receive the dividend diverges due to their payment willingness endowments.
The incremental AI value in the United States is absorbed by the existing subscription system. U.S. users are accustomed to paying high prices for software subscriptions, and cutting-edge model manufacturers compete with existing software giants for the same ecological niche. Microsoft bundles AI functions into high-priced subscriptions, converting one-time capability advantages into continuous customer payments.
The Chinese market is constrained by low willingness to pay for software subscriptions. For the same basic office software, the pricing in the Chinese market is usually only 1/5 or even lower than that in the United States. Therefore, midstream manufacturers in China generally price according to computing power cost, and the value of low-price Tokens directly overflows to the application layer.
For Chinese application layer companies, the migration of pricing methods determines the direction of profits. If subscription pricing is maintained, the Token price reduction dividend remains on the cost side, reflected in the improvement of gross profit margin and operating leverage; if pricing shifts to Token consumption or business results, the dividend will directly enter the revenue side.
However, the change of pricing method does not mean that profits are secured. Whether the downstream can retain profits depends on the "migration cost" established in the scenario, which is the only weapon to resist the buyer's price suppression and dividend recovery.
The attributability of output value, the exclusivity of scenario assets, customer structure and the consumption intensity of tasks determine the quality of scenarios. The moat within the scenario is maintained by three types of carriers: first, access qualifications and procurement systems granted by external rules (such as government affairs and regulated industries); second, data docking and caliber precipitation accumulated over time (such as financial data governance and intelligent operation and maintenance); third, the reset cost brought by system-level integration.
With the collapse of the old seat-based billing system, the pricing model based on Token and effect will greatly shorten the verification cycle of customer stickiness. Customer retention, which used to be known only after the annual renewal, can now be clearly observed through the quarterly Token usage. The downstream players who can finally retain profits are those who bind customers to their own business flow with extremely high replacement costs.
This article is from the WeChat official account "Hard AI", Author: Kozmon, Editor: Hard AI, Published with authorization from 36Kr.