HomeArticle

"The cheaper the token, the more expensive the bill" — The U.S. AI budget landscape has split into three tiers, why are they collectively flocking to Chinese models?

36氪的朋友们2026-10-10 21:32
U.S. enterprises have tiered AI budgets, and Chinese open-source AI has gained access to the U.S. small and medium-sized market.

Under the cost pressure of Tokens, the AI budgets of US enterprises are gradually forming a three-tier structure.

The first tier consists of giants building their own alternatives, the second tier sees mid-sized enterprises slashing budgets or downgrading service tiers, and the third tier is made up of small and medium-sized enterprises that directly turn to Chinese open-source weights. Behind this trend of stratification, leading US labs like OpenAI and Anthropic are losing both ends of the market: the largest enterprise clients are building their own tools, while the smallest enterprise clients are flocking to Chinese AI models.

Back at the start of this year, a term was trending in Silicon Valley: TokenMaxxing — the more Tokens you burn, the higher your productivity, and budgets are meant to be spent. However, this "Token Great Leap Forward" only lasted for half a year, and the situation has now completely shifted, with no one mentioning TokenMaxxing anymore. When the bills come due, companies of different sizes have adopted completely different countermeasures. As a result, US enterprises' AI procurement has clearly split into three tiers, opening up a distinct market space for Chinese AI.

Current Status: Tokens Are Cheaper, But Enterprises Can No Longer Afford to Burn Them

The most striking figure comes from Uber. According to foreign media reports, Uber's entire 2026 AI programming budget was exhausted in less than four months. The spending was so fast because employees actively used AI as required by Uber. In March this year, among Uber's 5,000-person engineering team, the penetration rate of Claude Code surged from 32% all the way to 84%, with a single engineer spending $500 to $2,000 on tokens per month.

A more extreme case is OpenClaw, an agent framework behind the "shrimp farming craze". Estimates show that a single instance running autonomously for a full day can consume API costs equivalent to $1,000 to $5,000, while users only pay $200 per month for the Claude Max plan; by early April this year, there were more than 135,000 OpenClaw instances running outside. Anthropic therefore quickly blocked this path in early April, requiring separate billing for the usage of such third-party frameworks, which increased some users' expenses by ten times or even dozens of times.

Sixteen days later, GitHub suspended new registrations for Copilot Pro and Pro+, and the reason given by Vice President of Product Joe Binder was equally straightforward: the computing power consumed by agent workflows often exceeds the cost users pay in a month, and the cost of a few requests can cover the price of the plan.

There is a trend that is widely misinterpreted. On the one hand, the unit price of Tokens is indeed falling, dropping by an order of magnitude roughly every 18 months, but on the other hand, the Token consumption per single task is rising faster than the price is falling. This leads to a result that surprises many people: the cheaper the model, the more expensive the bill becomes.

Gartner predicts that global AI expenditure will reach $2.5 trillion this year, a year-on-year increase of 69%; but at the same time, they predict that 25% of the 2026 AI budget will be deferred to 2027. "Deferring tiers" is a decent way of saying cutting budgets. On the other hand, Gartner also predicts that only 28% of AI infrastructure projects have fully delivered the original business case. In other words, nearly three-quarters of the money spent has not proven to be worth the investment.

Tier 1: Giants Are Self-Sufficient

The first tier consists of AI giant companies that have the ability to build their own models, for whom cost saving and promotion of their own products can be achieved simultaneously.

After Microsoft replaced Claude Code, it encouraged engineers to use its self-developed MAI-Code-1-Flash. This is a small model with about 5 billion parameters, released at this year's Build Conference and rolled out to all GitHub Copilot plans; the subsequent version MAI-Code-1.1-Flash even claims to reduce costs by another 73%.

Behind these moves, it is reported that Microsoft has cut its budget estimate for internal Anthropic usage by more than one-third, while the previous estimate was at least $1 billion per year.

Microsoft's position in this matter is actually quite delicate. In November 2025, they joined hands with NVIDIA to invest $15 billion in Anthropic, pushing the company's valuation to $350 billion, and also launched Claude on Azure for external sales. However, investing in Anthropic and cutting the internal Claude budget are not contradictory. For Microsoft, Claude is just inventory, while MAI is its own core asset.

Meta was the AI giant that went the craziest for TokenMaxxing before, even taking Token usage as an assessment indicator. Under the pressure of performance reviews, their employees once consumed 60.2 trillion tokens through Anthropic's tools within 30 days. At the peak, Meta internally estimated that it would spend $10 billion a year on Anthropic models.

But starting from the end of June, they restricted employees from using Claude Code and OpenAI's Codex, with the AI application group facing the strictest restrictions, and some tasks even requiring an approval process. Their alternatives are MetaCode and Muse Code. According to media reports, the number of internal Claude users at Meta has dropped from about 60,000 to 30,000, which is derived from external estimates of token consumption and expenditure.

Even after the cut, Meta's monthly Claude expenditure is still reported to be at the level of hundreds of millions of dollars. There is another reason for Meta to tighten the use of Claude — model distillation. If the code written with Claude is included in its own training data, Meta's model may absorb the features of the competing model along with it.

Tier 2: Mid-Sized Companies Can Only Cut Budgets

The second tier consists of mid-sized companies like Uber: with thousands of engineers, their bills are at the same level as giants, but they do not have their own cutting-edge models to switch to. Uber's CTO Praveen Neppalli Naga admitted in April this year: after the company deployed Claude Code to about 5,000 engineers, it burned through the entire year's AI budget in just four months, which caught him off guard.

Uber's senior management intended to popularize AI tools among engineers, and from this perspective, they achieved their goal. Statistics in May this year show that 95% of Uber engineers use AI tools every month. But the senior management did not consider the other side: the API cost generated by a single engineer per month is as high as $500 to $2,000.

Faced with the staggering Token budget, even a large enterprise like Uber with a market value of $150 billion has to urgently formulate strict hierarchical management to limit employees' usage traffic, counting every Token cost as carefully as they saved paper in the past.

This has become a universal consensus in the industry. Aaron Levie, CEO of cloud service company Box, said that he attended a dinner attended by CIOs of Fortune 500 companies and found that the most discussed topic among business leaders was not macroeconomic issues, but the Token cost of their enterprises.

Levie listed five methods that CIOs are trying to cut budgets: assigning tasks to different models according to workload, issuing agents with different capability levels according to user types, setting different spending caps for each team, requiring teams to demonstrate the necessity of AI according to usage scenarios, and some simply letting it go without restriction.

After hitting the wall at this tier, there are actually only three choices: cut the number of seats, downgrade to use cheaper models, or push the budget to the next fiscal year. The 25% AI budget deferral ratio from Gartner is mainly contributed by this tier.

Since there are cheap and easy-to-use Chinese open-source models on the market, why don't these mid-sized enterprises switch directly? Because US enterprises at this level are more concerned about data security compliance issues. They would rather use less than take regulatory risks. Once they are caught with flaws by regulatory authorities, the fine amount may far exceed the saved IT costs.

The US government's attitude towards Chinese open-source AI has never been clear, and how to regulate it or whether to regulate it at all is still an unknown. On the open-weight industry open letter in July, the signatories included NVIDIA, Microsoft, Meta, IBM, Dell and Palantir; Anthropic did not sign, and their CEO Dario Amodei opposed a one-size-fits-all ban and advocated targeted regulation. When the rules are pending, the legal teams of mid-sized enterprises usually choose to stay put.

Occasionally, big customers emerge at this tier. Airbnb chose Tongyi Qianwen instead of ChatGPT in its customer service scenario, which is currently the most prominent sample, and it also shows that once the attraction of price crosses a certain threshold, even the procurement process of listed companies cannot resist it.

Tier 3: A Large Number of Startups Turn to Chinese AI

The third tier consists of small companies and startups. They have neither compliance burdens nor self-development capabilities, and the only decision variable left is cost-effectiveness.

For an AI census of this tier, we can look at the data from the model routing platform OpenRouter. The peak weekly share of Chinese models on this platform has reached 46.4%, while US models account for 35.7%; in the ranking of single suppliers, DeepSeek ranks first with 17.6%, about 5.13 trillion tokens per week, Tongyi Qianwen accounts for 13.9%, and Anthropic accounts for 14.8%.

The price difference explains almost everything. Justin Summerville from OpenRouter said that open-source Chinese models are 60% to 90% cheaper than the flagship products of Anthropic and OpenAI; for specific models, as of June this year, DeepSeek V4 Flash charges $0.14 per million input tokens, while GPT-5.5 charges $5, with the overall price difference ranging from 4 times to 100 times.

It needs to be explained that this 46% does not mean that Chinese models have captured nearly half of the US enterprise market. The user structure of OpenRouter is naturally biased towards developers and small and medium-sized teams, and the traffic of large customers who sign enterprise contracts directly with Anthropic and OpenAI is not included in this pool.

But the existence of the three-tier structure has become an undisputed trend. Chinese models getting 46% share on OpenRouter shows that they are occupying not the US enterprise market, but the US startup market.

For leading US labs, the worse thing is the future: they are losing their next generation of customers. Startups that start with DeepSeek today are very likely to not switch back even after they grow up in three years.

OpenAI Cut Prices Twice in a Row to Expand Its Market Share

Faced with the simultaneous pressure from these three tiers of enterprise AI budgets, OpenAI's response was to cut prices, and it cut prices twice in a row.

In the round on July 30, GPT-5.6 Luna was directly reduced by 80%, to $0.20 per million input tokens and $1.20 per million output tokens, and Terra was reduced by 20% to $2 and $12; in the round on September 22, GPT-6 Sol was reduced to $2 and $10, while the previous generation GPT-5.6 Sol was $4 and $20, GPT-6 Luna was reduced to $0.10 and $0.50, and cached input reading was directly discounted by 90%.

OpenAI clearly stated that this is the permanent price rather than a promotion. The explanations OpenAI gave for the price cut all point to internal efficiency: hardware routing optimization, inference software improvement, context caching improvement, and they also mentioned that Sol autonomously rewrote the production kernel in supervised tests, cutting service overhead by 20% and increasing token generation efficiency by more than 15%.

Although they did not mention competitors, the direct cause of these two rounds of price cuts has a very clear market positioning. Anthropic's Claude Opus 5.5 went online about 90 minutes before OpenAI's September release, priced at $4 and $20, while GPT-6 Sol is exactly half of that. OpenAI CEO Sam Altman said on X that GPT-5.6 Sol is half the price of Claude Fable 5, and he is "happy to deliver at a quarter of the price".

Of course, the continuous upward attack of Chinese models in the low-end market is a real threat. One detail is very thought-provoking: the indicator that OpenAI is starting to emphasize now is "intelligence per dollar". Faced with the price gap of the order of magnitude between $0.14 for Chinese AI and $5 for US AI, US manufacturers have given up direct price comparison and can only emphasize their own intelligent performance.

Under the cost pressure of Tokens, US enterprises' AI budgets present a three-tier structure. The first tier builds their own alternatives, the second tier cuts budgets or downgrades service tiers, and the third tier directly turns to Chinese open-source weights. Leading US labs like OpenAI and Anthropic are losing both ends of the market: the largest enterprise clients are building their own tools, while the smallest enterprise clients are flocking to Chinese AI models.

This article is from the WeChat Official Account "Sina Tech", written by Zheng Jun, and published by 36Kr with authorization.