The Revelations of the Musk vs. Kimi Spat: Large Language Models Have No Moat, But Model-Building Enterprises Do
The competition will focus on model ecosystems rather than duels between individual models.
The person who sparked Musk's competitive instinct is not Liang Wenfeng from DeepSeek, but Yang Zhilin from Kimi, a 34-year-old young man from Shantou, Guangdong, who graduated from Tsinghua University and Carnegie Mellon University.
Recently, during the 2026 World Artificial Intelligence Conference, Kimi launched its "model nuclear bomb" — Kimi K3, with a parameter count as high as 2.8 trillion. After commenting that it was "Impressive", Musk announced that his company's new 2-trillion-parameter model would complete initial training next week and "may surpass Kimi". Kimi responded: Welcome to the 2-trillion club.
In the past two years, China witnessed the "hundred-model battle", and now it has evolved into a peak showdown across the two sides of the Pacific. The extremely fast iteration speed of large models brings new insights to observers of the AI industry: large models themselves are being commoditized at an unprecedented speed, but the competitive barriers of large model enterprises are shifting from the model layer to dimensions that are harder to replicate.
The moat at the model layer is collapsing on a weekly basis
The parameter count of open-source models jumped from 1 trillion to 2.8 trillion in just two weeks. In contrast, when ChatGPT-2 was released in 2019, its parameter scale was only 1.61 billion.
The table below shows that the parameter inflation of open-source models is accelerating increasingly.
This sense of speed reflects that in the AI industry, the "moat" built by parameter scale is shortening on a weekly basis.
Behind this is the sense of urgency brought by the fierce competition among large models. Everyone is aware that: The window for large models to seize their positions is extremely short.
Back in 2025, consulting firm Gartner released the "2025 Hype Cycle for Artificial Intelligence Technologies", making a harsh prediction that the market is entering a contraction phase and will form a "core-satellite" pattern, where only 20% of model suppliers can survive beyond 2026.
The base models at the bottom of the pyramid will be commoditized (i.e., as ubiquitous and essential as water and electricity), and the competition focus will shift to the "Agentic AI" at the top of the pyramid.
Li Kaifu, founder of 01.AI, also pointed out in March 2025 that there are only 5-6 truly competitive base models worldwide (such as GPT, Claude, Gemini, and 1-2 from China), and eventually only three AI model companies will remain in the Chinese market: Alibaba, ByteDance, and DeepSeek.
Now it seems that these predictions are generally in line with the evolution direction of the industry.
At the beginning of 2025, the cutting-edge models of OpenAI and Anthropic were clearly ahead of any self-deployable models. But by mid-2026, this gap has been narrowed to just a few percentage points in most benchmark tests.
From this macro perspective, looking at the release of Kimi K3, we can see that it is striving to reach the top tier and secure a key position. Through the open-source approach, anyone can download, deploy, and fine-tune it. With its massive parameters, it aims to become a top-tier infrastructure as essential as electricity.
Its strategic goal is to seize the commanding heights of the market, so that all future revenue generated within the ecosystem will be based on this infrastructure.
Because as parasites on top of the model, players generally won't ask "which power plant provides better electricity". Instead, they will only ask "which power grid is more stable, cheaper, and has wider coverage".
Where has the real moat gone?
If the model itself is not the moat, then where is the real barrier?
A growing number of industry insiders believe that we should not judge players solely by their models. The barriers of models lie in three dimensions that are harder to replicate and harder to overcome in the short term: computing power and cost structure (underlying barrier), distribution channels and ecosystem binding (mid-tier barrier), and trust, compliance, and scenario binding (application-layer barrier).
The first layer: Computing power and cost structure (underlying barrier) — competing to become smarter with less spending.
Gabe Goodhart, Chief Architect of IBM AI Open Innovation, put it very bluntly in an interview earlier this year: "The model itself will not become the main differentiator."
He believes that 2026 will usher in the homogenization stage of AI products, and "orchestration capability" — the ability to integrate models, tools, and workflows — will become the key to competition.
The reason behind this judgment is his conviction that computing power resources are evolving from an "optional asset" to a "lifeline". Since the beginning of this year, enterprises that control land, electricity, data centers, and advanced chip supply chains have gained favor in the capital market. Under this rigid constraint, whoever can optimize the combination of computing power, cost structure, and work efficiency will gain more initiative.
In this macro environment, we can see that the architectural design of K3 is strategically oriented to deliver "Chinese-style cost-effectiveness".
It adopts the Kimi Delta Attention (KDA) algorithm and attention residual mechanism. Each time processing a token, it only activates 16 out of 896 expert modules, with merely 1.8% of the experts participating in the actual computation.
Kimi claims that this architecture improves the overall scaling efficiency by about 2.5 times compared to its predecessor Kimi K2.
Its technical investment direction is not to "pile up larger parameters", but to "achieve better performance at lower costs", striving to optimize both cost and effect in the same direction.
The second layer: Distribution channels and ecosystem binding (mid-tier barrier), which essentially aims to attract more developers and form a natural traffic generation mechanism.
OpenAI's path is to build distribution barriers by leveraging Microsoft's ecosystem. GPT-series models are deeply embedded in products like Microsoft 365 Copilot, similar to a "parasitic distribution" approach, reaching hundreds of millions of enterprise users worldwide through the existing giant.
Anthropic, on the other hand, breaks through by focusing on the enterprise market and heavily investing in programming capabilities. According to the Ramp AI Index data in April 2026, Anthropic's share in U.S. enterprise AI spending has reached 34.4%, surpassing OpenAI's 32.3% for the first time.
In contrast, Kimi regards open source as a customer acquisition tool, accumulating its own ecosystem through developers.
In February 2026, programming tool OpenClaw announced that it would set Kimi K2.5 as its official primary model, which is regarded as a key turning point for Kimi's counterattack.
Rather than burning money on user growth, letting the developer ecosystem "bring its own traffic" is another form of parasitic model. Kimi is clearly striving to minimize the cost of building the ecosystem.
These three paths ultimately converge to the same goal: when the model itself cannot form barriers, "being widely used" is more important than "being technically excellent".
The third layer: Trust, compliance, and scenario binding (application-layer barrier), which builds the soft power of the model.
Against the backdrop of increasingly strict global AI regulation, the importance of this layer is rising rapidly. Many companies may have overlooked compliance costs, which are actually fixed expenses. Large companies undoubtedly have advantages here, while small companies find it difficult to amortize these costs.
Just look at Anthropic: starting from its stance on AI safety, it is turning this potential business disadvantage into a competitive advantage, expanding its territory in fields such as healthcare, law, finance, and government.
In Silicon Valley, this is regarded as a "moral premium". In Washington, it is seen as "auditability". No matter how it is interpreted, it has translated into tangible contracts and revenue.
Kimi's strategy of pursuing both
The release of K3 by Kimi, to some extent, has torn off the fig leaf of the industry — the model itself barely has any moat. But at the same time, Kimi is also an explorer of "model enterprises with moats" — it aims to survive through channels and speed.
Now it seems that Kimi's open source is a business strategy to build moats, not a technical idealism.
Kimi has settled in mainstream cloud platforms and integrated into developer toolchains, completing global distribution through a "parasitic" rather than "self-built" approach. While others are still arguing about "whether open source will weaken competitiveness", Kimi has already exchanged open source for free promotion from developers worldwide.
Once it has a developer base, Kimi does not hesitate to push forward commercialization.
The API pricing of K3 is $3 per million tokens for input and $15 per million tokens for output. Observers consider this a fairly aggressive pricing. The dual-track strategy of open-source weights + commercial APIs allows Kimi to gain both widespread recognition and substantial revenue.
In fact, Kimi never fights an unprepared battle. It has already tested the feasibility of this dual-track strategy. After the release of Kimi K2.5 and the viral success of Kimi Claw, Moonshot AI's cumulative revenue in 20 days exceeded its total revenue for the whole of 2025.
According to Stripe data, the number of Kimi personal subscription payment orders surged by 8280% month-on-month in January 2026, and rose by another 123.8% in February. Its ranking on Stripe's global list jumped from outside the top 100 to the top 10, making it the first Chinese AGI product to enter the top 10 of this list.
With ARR breaking through $100 million in March, doubling to over $200 million in April, and exceeding $300 million by mid-June, this growth rate itself is a barrier.
The capital market is willing to continue providing funding, which allows the company to continuously purchase computing power, recruit talents, and expand channels, forming a positive feedback loop.
Along with the rapid progress in commercialization, Kimi's capital moat is also getting wider. Its valuation rose from $4.3 billion to $31.5 billion within half a year, with an increase of over 700%. This sets a huge threshold for all followers.
However, Kimi is not completely without risks.
On July 19, Moonshot AI released an announcement: within 48 hours after the release of K3, user requests far exceeded expectations, approaching the carrying limit of the existing cluster. Starting from that day, new C-end user subscriptions were suspended, which can be described as a sweet trouble.
Homogenized competition is also intensifying. In terms of long-context capabilities, Kimi and DeepSeek are both in the first echelon of domestic models, and Kimi does not have exclusive advantages, which indicates that the moat between domestic players is also very shallow.
Overseas, the United States is considering restricting local enterprises from adopting Chinese AI models. If implemented, it will directly impact Kimi's newly established overseas distribution channels.
An open question without a conclusion
Kimi's K3 strategy seems impressive, but it is not the only path for Chinese large models to build moats.
Zhipu AI, which competes against Kimi, did not take the cost-effectiveness path of "reducing model costs", but chose the path of "raising prices". Since the first quarter of this year, Zhipu AI's API pricing has increased by 83% month-on-month, but the call volume has not decreased; instead, it has risen by 400%.
This shows that once customers form workflow dependencies — such as prompt habits, internal toolchains, and fine-tuning data — their price elasticity will be very low. This is essentially a soft moat formed by switching costs, rather than a moat brought by technological leadership.
In this way, Zhipu AI forms a sharp contrast with Kimi: the former seizes the market through "price increase + user locking", while the latter seizes the market through "open source + speed".
But both cases illustrate one thing: the moat of large models no longer lies in the model itself, but in the intertwined user habits, ecosystem locking, and cost structure built by large model enterprises.
From an industry perspective, when the model layer is fully commoditized, the computing power layer is monopolized by giants, and the distribution layer is locked by ecosystems, what remaining cards do small and medium-sized AI companies have to play?
OpenAI CEO Sam Altman once said that the future large model market will present a pattern of "a handful of top-tier models + countless downstream callers". Is this prediction really going to come true?
Overall, the upcoming battle among models will become a position-seizure war among giants, with an unprecedented level of intensity. But no one can guarantee that they will be the eternal king.
Musk's sharp attitude shift from "liking" to "declaring war" to "possibly surpassing" precisely shows that even the top players in the world cannot guarantee absolute victory.
This exposes an awkward fact: in an era where model capabilities are rapidly converging, even the definition of "surpassing" itself has become ambiguous.
Kimi's remark "Welcome to the 2-trillion+ club" appears to be confident on the surface, but reflects deep sobriety underneath. It knows that 2.8 trillion parameters will never be a permanent barrier. The real game lies beyond parameters.
This article is from the WeChat official account "Muhe's Tech Launch Event", written by Gong Zheng, and published by 36Kr with authorization.