HomeArticle

Large AI models have entered the "Infinity War"

袁振2026-07-23 12:33
The era without moats is precisely the period when the competitive landscape is yet to take shape.

In the early hours of July 17, 2026, Elon Musk typed a single word on X: "Impressive."

He was replying to an Artificial Analysis review post about Kimi K3. This new model, released by Moonshot AI, has a parameter scale of 2.8 trillion, making it the world's largest open-source weight model. In the Artificial Analysis Intelligence Index, K3 ranked third globally at launch, behind only Claude Fable 5 and GPT-5.6 Sol[1].

The next day, Musk added a comment while talking about the 2-trillion-parameter model that xAI is training — it "could surpass Kimi". Moonshot AI responded on Weibo: "Welcome to the '2 Trillion+' club"[2].

But what truly shook the market was not the ranking on the leaderboard.

On the evening of July 19, Kimi released an announcement: Since the launch of K3, "user requests over the past 48 hours have far exceeded expectations and are approaching the load limit of the existing cluster". The company decided to suspend new C-end user subscriptions immediately and allocate all existing computing power to support subscribed users[3].

A model that crashed itself just two days after launch — the last time this level of attention was seen was with DeepSeek R1 in January 2025.

And in the same week that K3 went viral, news broke that the official version of DeepSeek V4 had started its gray-scale testing, with a scheduled release at the end of July. When the preview version was launched in April, V4 Pro had 1.6 trillion parameters, with an API input price of $0.435 and an output price of $0.87 per million tokens[4]. Artificial Analysis' standardized evaluation shows that K3 spends an average of $0.94 on a weighted evaluation task, while V4 Pro costs about $0.04 — a difference of roughly 23.5 times[2].

If we rewind the clock further —

During the 2026 Chinese New Year, ByteDance, Alibaba, and Tencent collectively poured tens of billions of yuan into AI products. ByteDance secured a partnership with the Spring Festival Gala, with Doubao AI interactions reaching 1.9 billion times on New Year's Eve; Alibaba spent 30 billion yuan on the Qwen "Treat Everyone" campaign, pushing Qwen's DAU to a peak of 73.52 million, nearly matching Doubao's 78.71 million; Tencent invested 1 billion yuan, and Yuanbao's daily active users exceeded 50 million[5][6].

This is not a single-point battle. It is a war without boundaries.

A Card Table Overturned

To understand what is happening today, we must go back to January 20, 2025.

On that day, DeepSeek released R1. The model is fully open-source under the MIT license, with API pricing: $0.14 per million tokens for input on cache hits, $0.55 on misses, and $2.19 for output. It also open-sourced 6 distilled versions ranging from 1.5B to 70B[7].

A week later, on January 27, NVIDIA's stock price plummeted nearly 17%, with its market value evaporating by about $589 billion in a single day — setting the record for the largest single-day market value loss of a single stock in US stock history at that time[8].

What the market feared was not how powerful a Chinese company's model was. It feared that R1 proved one thing: near-cutting-edge capabilities can be obtained at low cost, downloaded instantly, and compressed into models of different sizes.

This is not selling models — it is rewriting the pricing system.

Back in May 2024, DeepSeek released V2, with API pricing of 1 yuan per million tokens for input and 2 yuan for output, which was nearly 1% of GPT-4 Turbo's price. This directly triggered a domestic large model price war: Zhipu AI took the lead by cutting prices by 80%, ByteDance's Doubao offered an ultra-low price of 0.0008 yuan per thousand tokens, Alibaba's Tongyi Qwen then cut prices by 97%, and Baidu and Tencent even announced that some models would be free of charge[9].

In an interview with Waves in July 2024, DeepSeek founder Liang Wenfeng roughly said a passage that still reads calm and sharp today:

"We never intended to be disruptors — all of this just happened by accident. We just followed our own pace, calculated costs, and priced reasonably. Our principle is not to sell at a loss, nor to pursue excessive profits." [10]

The power of this statement lies in: While others were burning money to grab market share, DeepSeek claimed it could be profitable at the cost line. While others priced models as scarce goods, DeepSeek treated them as supply goods.

The essence of the "DeepSeek Moment" is not technological leadership. It is three things happening simultaneously: capabilities reaching the frontier, prices plummeting, and weights and distilled models being instantly available. These three things combined turn "powerful models" from weapons for the few into tools for everyone.

In 1950, Claude Shannon, the father of information theory, published a paper titled Programming a Computer for Playing Chess. He chose chess as the starting point for computer games for a simple reason — "the problem is clearly defined, operations are clear, and goals are determined". This paper was later regarded as a foundational work in the field of computer games (although the term "artificial intelligence" would not be formally proposed until the Dartmouth Conference six years later). Shannon could hardly have imagined that the "well-defined toy problem" he chose back then would grow into a $2 trillion industry 75 years later, where even pricing power remains unclear.

K3's Bet: No Low Prices, Long-Term Play

K3 took a completely different path.

Technically, K3 is built on a self-developed KDA hybrid linear attention mechanism and Attention Residuals, activating only 16 out of 896 experts at a time, paired with a 1 million-token context window. Officials say these structural improvements have increased K3's overall scaling efficiency by about 2.5 times compared to K2[1].

In terms of pricing, K3 API costs 20 yuan per million tokens for input (2 yuan on cache hits) and 100 yuan per million tokens for output. This is not in the same price range as DeepSeek V4 Pro[1].

But K3's logic is not selling tokens — it is selling "task completion".

At the Zhongguancun Forum this March, Yang Zhilin, founder of Moonshot AI, systematically elaborated on his three scaling directions: improving token efficiency, extending context, and organizing Agent clusters. His first principle is simple — "The essence of building large models is to convert more energy into intelligence"[11].

In less than four months, K3 became almost the physical embodiment of these three judgments. The Kimi Work desktop client, launched in June, can read local files, access browsers, write code, and deliver documents. K3 has been integrated into Kimi.com, Kimi Work, Kimi Code, and APIs simultaneously[2].

Yang Zhilin is not betting on "my model is smarter than yours" — he is betting on "my model can work for you".

In 1990, Steve Jobs said in an interview that was lost for many years and not discovered until after his death: "Computers are bicycles for the mind." What he meant was that tools do not replace humans — they only amplify human capabilities, just like a bicycle does not pedal for you, but makes you pedal faster. But K3's logic is: I will not be your bicycle — I will pedal for you directly. Jobs probably would not object, after all, he also said "the people who are crazy enough to think they can change the world are the ones who do". But this time, it is a machine that is crazy.

This represents a fundamental paradigm shift. Over the past two years, the competitive logic of large models has been "I answer better than you" — competing on benchmarks and leaderboards. But K3's logic is "I can get the task done" — competing on who can complete a task from start to finish without errors, deviations, and with deliverables.

But this path has a core problem.

Rich Privorotsky, a partner at Goldman Sachs, warned after K3's release: "The era of computing power expansion may have come to an end, and AI valuation logic is facing reshaping." He pointed out that what is truly thought-provoking is how a Chinese lab that cannot match the Western world's largest pre-training computing power has rapidly narrowed the gap with top US models through architectural innovation. "Scaling is no longer the only winning path."[3]

And for K3, the challenges are more concrete. The computing power outage on July 19 has shown that the inference cost of long-thread tasks is far higher than that of chat interactions. The higher the popularity, the greater the cost pressure. If scaling up only brings more expensive inference, the "long thread" will turn into a growing cost sheet.

Giants Enter the Arena: Models Are Weapons, Entrance Is the Target

If you only look at DeepSeek and Kimi, you might think the large model war is between "technological idealism vs. task executionism". But bringing in ByteDance, Alibaba, and Tencent paints a completely different picture.

These three are not selling models. They are using models to seize entrances.

The 2026 Chinese New Year was the full outbreak of this entrance war.

ByteDance tied itself to the CCTV Spring Festival Gala, with Doubao AI's total interactions reaching 1.9 billion times on New Year's Eve alone, generating over 50 million New Year avatars and over 1 billion New Year greeting messages[5]. ByteDance's logic is clear: divert traffic from the Douyin pool, acquire new users through national-level exposure at the Spring Festival Gala, and retain users with an "emotional companionship" positioning.

Alibaba took a different path — letting AI directly "treat everyone". Spending 30 billion yuan, users can say "help me" to Qwen to buy milk tea for 1 cent, get movie tickets with red envelope subsidies, and book flights and hotels. Official data shows that during the Spring Festival, users placed nearly 200 million "one-sentence orders" on Qwen, of which more than 4 million people aged 60 and above completed their first food delivery order in life through AI[5].

Alibaba's plan is: turn AI from a chat tool into a life entrance. When users get used to saying "help me buy a cup of milk tea" to Qwen, Qwen will no longer be just a model — but a super hub connecting Alibaba's ecosystem (food delivery, ticketing, payments).

Tencent's Yuanbao, on the other hand, bets on "AI + social". Investing 1 billion yuan in red envelopes, it launched the group chat AI feature "Yuanbao Pai", allowing users to invite WeChat friends to join and complete various tasks by @-mentioning Yuanbao in group chats. Ma Huateng hopes that Yuanbao can replicate the success of WeChat Red Packet back then[5][6]. During the Spring Festival, Yuanbao's daily active users exceeded 50 million, and its monthly active users reached 114 million[6].

But the side effects of the red envelope war are also obvious. According to TingTing Tech, many young people in county towns uninstalled the app after the campaign ended, "They know AI can draw pictures and chat, but they can't figure out what this has to do with their daily lives"[5].

Subsidies can buy downloads, but not usage habits. This is a lesson that has been proven countless times in the Internet era, and it still applies in the AI era.

Five Paths, Five Bets

Lay the cards of the five companies together, and you will find an interesting structure:

DeepSeek's logic is "make intelligence cheap". Liang Wenfeng put it plainly in the interview: "We believe that AI and API services should be affordable and accessible to everyone."[10] It bets that if intelligence is cheap enough, the application ecosystem will grow naturally. The risk is that the price war will continue to compress profit margins.

Kimi's logic is "make intelligence capable". Yang Zhilin bets that users are willing to pay for a higher probability of task completion. If a model can compress an engineer's one-week work into one day, token cost will no longer be an issue. The risk is that user and enterprise procurement systems are priced by tokens, and changing the billing unit takes time.

ByteDance's logic is "make intelligence companionship". Doubao takes an emotional route, keeping users engaged because they "want to chat" rather than "have to use it". It has Douyin's traffic pool, but traffic does not equal retention.

Tencent's logic is "make intelligence scenarios". Yuanbao is connected to JD E-commerce and Meituan Food Delivery, embedding AI into real business closed loops through the WeChat ecosystem. It has the most scenarios, but if model capabilities lag behind for a long time, the scenarios may be "parasitized" by stronger models.

Alibaba's logic is "make intelligence infrastructure". Qwen takes a full-stack route, integrating bottom-layer models, middle-layer cloud services, and upper-layer applications. It competes with DeepSeek for developers, with ByteDance for users, and with Tencent for scenarios. Its front line is the longest.

Five paths, five assumptions, five different bets.

But the most fundamental question of this war is not in the hands of any company. It lies in a more underlying question: What on earth should the "intelligence" of large models be priced by?

By token? That's DeepSeek's logic. By task? That's Kimi's logic. By user duration? That's ByteDance's logic. By transaction closed loop? That's Tencent's logic. By cloud consumption? That's Alibaba's logic.

In the 1990s, Microsoft proved with Windows that the pricing power of operating systems belongs to the "platform". In the 2000s, Google proved with Search that the pricing power of information distribution belongs to the "entrance". In the 2010s, WeChat proved that the pricing power of user relationships belongs to the "ecosystem".

In the 2020s, who will decide how much one unit of "intelligence" is worth?

Wittgenstein wrote in Tractatus Logico-Philosophicus: "The limits of my language mean the limits of my world." The "language" he referred to was logical language, not large models. But today, this sentence suddenly takes on a completely different meaning: large models are expanding the "limits of language" at a speed humans cannot catch up with — from chatting to writing code, from writing code to calling tools, from calling tools to completing tasks independently. Every expansion of the boundary is redefining what "intelligence" can do and how much it should cost. Wittgenstein could hardly have imagined that his philosophical contemplation on the relationship between language and the world would be verified 75 years later by a group of GPU builders in the most brutal way — through a price war.

In 1950, the same year Shannon wrote his chess paper, Turing asked a question