HomeArticle

Major AI players have waged the "Token Milk Tea War"

字母AI2026-08-01 14:25
Models are becoming more and more similar, and those that gain solid user retention will dominate the entire market.

Over the past two months, AI companies have suddenly started issuing "coupons" en masse.

On July 31 Beijing time, OpenAI announced that it will cut the API price of GPT-5.6 Luna by 80%, reduce the price of Terra by 20%, and the quota consumed by calling these two models in Codex and ChatGPT Work will also decrease accordingly.

For a moment, Tokens gave people the exact same impression as discount coupons. They can be either new product trial coupons, or compensation coupons after system failures. They can be packed into membership plans, or designed as "medium size, large size, extra large size", allowing users to refill after the quota is used up.

This is very similar to the "milk tea war" that just broke out in the food delivery industry: platforms compete to issue coupons, which seems to benefit users, but in essence they are fighting for users' consumption habits.

In the past, competition among AI companies mainly focused on parameters, ranking lists and model launch events. However, with the explosive growth of Agent users, narrowing gaps between models and declining migration costs, technical capabilities have only become an admission ticket to the market.

Nowadays, AI companies need to face a more direct problem — when users can switch between different AI tools at any time, how to make them stay?

AI companies have started "issuing coupons"

In this round of Token promotions, a landmark signal appeared around May 20.

At that time, OpenAI proposed to YC startups that each company could get up to $2 million worth of Tokens, on the condition that they give up a portion of their equity.

OpenAI is essentially using its own computing power to participate in investment, which not only places bets on these startups, but also hopes that they will use OpenAI's models from the very beginning and continue to become paying customers in the future.

In June, the use of Tokens became even more similar to e-commerce coupons. After Codex over-deducted user quotas due to system anomalies, OpenAI uniformly reset user quotas twice, and additionally gave away one extra reset opportunity that could be saved for later use.

When something goes wrong with the system, OpenAI compensates users with quotas, which is equivalent to giving out a "refill coupon".

Anthropic took a different path.

Fable 5 was initially only free for one week, but the event was extended twice in a row, and the activity of increasing Claude Code's weekly quota by 50% was also extended simultaneously. Starting from July 20, Fable 5 was officially included in the premium plan, and the free experience originally used for new product promotion eventually became a long-term membership benefit.

DeepSeek's move was even more straightforward.

V4-Pro first cut its price permanently by 75% to attract developers with low prices; at the end of June, it proposed to adopt peak-valley pricing to control usage through time-of-use pricing. MiniMax packaged Tokens into three tiers of plans priced at $20, $50 and $120 respectively, corresponding to different usage volumes, which looks exactly like the medium, large and extra large sizes on a milk tea shop menu.

Now OpenAI has also directly joined the price reduction wave. On the last day of July, OpenAI announced that both the API input and output prices of GPT-5.6 Luna have been cut by 80%, now standing at $0.2 and $1.2 per million Tokens respectively; Terra's price was reduced by 20%, down to $2 and $12 respectively. The subscription price and total quota of Codex and ChatGPT Work remain unchanged, but less quota will be deducted when using Luna and Terra. The flagship model Sol keeps its original price.

This "Token milk tea war" is not only happening in membership plans. For high-value enterprise customers, manufacturers give away benefits even more directly. According to reports from The Wall Street Journal, the CEO of AI customer service platform Pylon revealed that since the beginning of this year, the company has obtained free Tokens worth about $1.6 million from one model supplier, and the other two have also given away tens of thousands of dollars worth of quotas respectively, and some manufacturers even provide several months of unlimited free usage rights.

Model price cuts and quota giveaways are nothing new.

But in the past, such activities were usually concentrated in the new product launch period, or used as developer support measures, and would end after a short period of excitement. However, in the past two months, the ways of giving away Tokens have emerged in an endless stream. Tokens not only started to function like coupons, being systematically used for user acquisition, retention and compensation, but also became bidding chips for competing for enterprise customers.

No matter how manufacturers package this behavior, it all points to the same fact: AI companies are no longer satisfied with pay-as-you-go pricing, and have begun to compete for users in a more "Internet-style" way.

Behind this is the fact that the single subscription system is increasingly unable to cover the vastly different usage intensities of users.

Light users value low entry barriers, while heavy users suffer from quota anxiety when performing long tasks. As a result, membership tiers, traffic packages, Credits and peak-valley pricing all emerged together. AI companies are no longer just selling models, but have also begun to study the business that consumer giants are very familiar with — how to get users to place their first order, what methods to use to make them buy one more cup afterwards, and what reason to give them to come back after they finish drinking.

The battlefield for AI companies has not left technical competition, but there is clearly an extra layer of commercial competition.

Or to put it in a more Internet-style way: the one who wins user retention wins the market.

Models are getting more and more similar, users are getting more and more fickle

Behind the manufacturers' sudden increase in investment in Tokens are two changes. The number of Agent users is increasing rapidly, and the quota consumed by a single user is getting higher and higher; the gaps between leading models in many common tasks are also narrowing, making it easier for users to switch platforms than before. Prices, quotas and membership benefits have therefore become tools to compete for users' usage habits.

First of all, AI users are still pouring in rapidly.

By the beginning of June, Codex had more than 5 million weekly active users, which was more than 6 times the number when its desktop application was launched in February.

More notably, non-developers account for about 20% of them, with a growth rate more than 3 times that of developers. The use of programming Agents has gone beyond writing code, and organizing materials, writing documents and analyzing data have also become common use cases. Some people even hand over the entire workflow of tasks to it for completion.

The way users consume Tokens has also changed. According to official OpenAI data, by May this year, 80.6% of Codex individual users have submitted at least one task equivalent to more than 30 minutes of human work, 25.6% have submitted tasks equivalent to more than 8 hours of work, and more than 10% of users have managed more than 3 Agents per week at the same time.

In the past, the process of a user asking a chatbot a question and getting an answer was simple, and the Token consumption was relatively small. Now, an Agent task may last for several hours. Token consumption has suddenly moved from the "2G era" to the "5G era", and the previous quotas are obviously no longer sufficient.

Anthropic's analysis of about 400,000 Claude Code sessions also shows that users use it for about 20 hours per week on average. From October 2025 to April 2026, the proportion of sessions using it to write documents and analyze data doubled from about 10% to 20%, and the average task value increased by 27%.

As Agents move from occasional trial use to formal work processes, quotas have changed from a bonus benefit to a rigid demand.

Manufacturers have begun to compete for users with high usage frequency and long payment cycles. By February this year, Claude Code's annualized revenue had exceeded $2.5 billion, doubling from the beginning of the year, the number of enterprise subscriptions increased by 4 times, and enterprise customers contributed more than half of the total revenue.

Whoever can make this group of users use Agents more frequently will likely obtain continuous growth in Token revenue.

The problem is that these users are not loyal at all.

The CEO of AI customer service platform Pylon put it very directly: "I can't see any loyalty at all."

The era of buying traditional software and sticking to one vendor for continuous payment is over. Enterprises now usually access several models at the same time, and which model to assign to a specific task mainly depends on performance and price.

Cursor's product itself is "model-agnostic", allowing users to switch between OpenAI, Anthropic, Google and open source models. Its person in charge said that in the past, the best model for a task might only change once every few months, but now it may be changed several times a week.

This also explains why manufacturers are in a hurry to give away more Tokens.

Agent users are growing rapidly, but most people's usage habits have not yet taken shape, and enterprises have not decided which model to integrate into their business in the end. Manufacturers are competing to become the first choice tool for individual users, and the default model in enterprise workflows.

In 2026, researchers from the University of Trieste, King's College London and University College London analyzed 7156 GitHub pull requests generated by programming Agents, and the results showed that no single tool can lead in all tasks. In most cases, the gap brought by task types is larger than the gap between different tools.

For users, the best model is increasingly becoming a dynamic multiple-choice question.

Migration is also getting easier.

DeepSeek provides both OpenAI-compatible and Anthropic-compatible interfaces, and MiniMax's Token plans can be directly connected to Claude Code, Cline and OpenClaw.

Developers don't even need to change their original working interface, and they can transfer tasks to other models just by replacing the underlying model.

AI companies are willing to provide Token subsidies to compete for users' long-term workflows. Once a model is integrated into an enterprise's system, the migration cost will rise rapidly, and the moat of AI products will be established. Therefore, early subsidies have obvious significance for seizing entry points.

Model capabilities are responsible for attracting users, while subsidies are responsible for giving users a reason to "not leave yet".

Technology is only an admission ticket,

AI companies are starting to catch up on commercial courses

But there is a fundamental difference between Token discounts and "milk tea coupons": giving one extra cup of milk tea has relatively clear and controllable costs. But giving away an extra 100 million Tokens, it is very difficult to calculate in advance how much computing power expenditure it will eventually translate into.

And computing power is still very, very precious at the moment.

As a result, an interesting scene has emerged — AI companies are issuing coupons on one hand, while limiting traffic on the other.

On May 6, Anthropic announced a computing power cooperation with SpaceX, which allowed Claude Code's 5-hour quota to double, and the limit would no longer be lowered during peak hours. After Fable 5 was restored, it can only account for up to 50% of paid users' weekly quota, and users need to purchase Credits for the excess part. The free experience has drawn a clear cost boundary from the very beginning.

MiniMax's large plans are also not for unlimited use. It sets 5-hour quota, weekly quota, maximum Agent concurrency and peak traffic limit at the same time, and suggests that users use pay-as-you-go pricing in production environments.

OpenAI also retains a clear cost boundary in this price cut. The subscription price and total quota of ChatGPT and Codex remain unchanged, only the quota deducted when using Luna and Terra is reduced. The price of the flagship model Sol remains unchanged, and the newly launched Fast mode of API provides up to 2.5 times the speed at twice the standard price.

Multiple Agent products of OpenAI still share a common quota pool, and complex tasks consume more quota. After the quota is used up, users can only purchase Credits, reset the quota, upgrade their plans, or wait for the next cycle to restore the quota.