Kimi K3, the product that most closely resembles DeepSeek, fails to replicate the decisive knockout strike.
K3 has suddenly become the focus of widespread attention.
The first to spark the buzz was Elon Musk.
In the early hours of July 17, Beijing time, Musk replied "Impressive" under a Kimi K3 review post on Artificial Analysis. The next day, when talking about the 2-trillion-parameter model that xAI is training, he stated that it "will likely surpass Kimi".
Shortly after, Moonshot AI's Kimi responded on Weibo: "Welcome to join the '2 Trillion+' Club".
Screenshot from Weibo
Why the "2 Trillion+ Club"?
On the evening of July 16, Moonshot AI released its new-generation large model Kimi K3, where "2 Trillion+" refers to the parameter scale range of K3. According to Moonshot AI, Kimi K3 has a parameter scale of 2.8 trillion, a 1 million-token context window, native visual understanding support, a self-developed KDA attention algorithm, and activates 16 experts out of 896 for each token.
Discussions about Kimi K3 are not limited to overseas social media platforms.
Anastasios Angelopoulos, co-founder and CEO of Arena, said in a media interview that this might be the biggest model release of the year. Meanwhile, K3 took first place on Frontend Code Arena, and ranked third when the Artificial Analysis Intelligence Index was launched.
The shortage of computing resources further fueled K3's rapid rise to widespread attention.
On the evening of July 19, Kimi, the AI assistant under Moonshot AI, released the "Notice on Computing Resource Shortage and Suspension of New Membership Subscriptions", announcing that it will suspend new C-end user subscriptions effective immediately, allocate all existing computing resources to protect the rights of subscribed users, and push forward capacity expansion at full speed.
Amid the heated discussions, K3 has been labeled with topics like "the next DeepSeek" and "DeepSeek Moment 2.0". In fact, as of press time, the relevant license has not been disclosed, and the complete technical report has not yet been released.
However, if we compare the two directly, the differences immediately become apparent. Artificial Analysis measured that K3 spends an average of $0.94 to complete a weighted evaluation task, while DeepSeek V4 Pro costs about $0.04 under the same standard. K3 has entered the global frontier comparison, but it has not replicated DeepSeek's most critical breakthrough.
1
It does look like DeepSeek in 2025
The impact created by K3 mainly comes from the sudden change in its comparison benchmarks.
According to the announcement from Moonshot AI, K3 has 2.8 trillion total parameters, native visual capabilities, a 1 million-token context, and activates 16 experts out of 896 for each token. Artificial Analysis gave it a score of 57, ranking third at launch; in the real-world knowledge work evaluation GDPval-AA v2, it achieved 1668 Elo, higher than Claude Opus 4.8 and GPT-5.5, but lower than Claude Fable 5; in the AA-Briefcase long-term knowledge work evaluation, it ranked second only to Fable 5 at launch.
Screenshot from X
"Although its overall performance still lags behind the most powerful proprietary models Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrates frontier-level performance in our evaluation suite, consistently outperforming other tested models." The report "Kimi K3: Opening Frontier Intelligence" states.
All these rankings have their boundaries.
As of July 21, K3 has moved from third place to fourth place after its launch. And Frontend Code Arena measures the visual, interactive, and user preference performance of frontend pages, which does not equate to full code capabilities.
However, these changes cannot deny that Kimi has completed an identity transformation: K3 is no longer just "a model that performs well among domestic models", but has directly entered the global comparison system composed of Fable, GPT, and Opus.
An intuitive change is that public attention has shifted from K3 to Yang Zhilin himself.
After the release of K3, Yang Zhilin's doctoral supervisor Russ Salakhutdinov publicly congratulated him on X; many media outlets have published profiles of him, and overseas tech media even wrote a separate article about the naming convention of meeting rooms at Moonshot AI.
Screenshot from X
After graduating from Tsinghua University in 2015, he went to Carnegie Mellon University to pursue his doctorate, participating in research projects such as Transformer-XL and XLNet, until he graduated in 2019. Some media reported that during his doctoral studies at Carnegie Mellon University, Yang Zhilin was known for his love of rock bands like Pink Floyd. This personal interest later influenced the company's naming: "Moonshot AI" is derived from Pink Floyd's album of the same name, and Kimi is Yang Zhilin's English name.
Pink Floyd released an album titled *The Dark Side of the Moon* in 1973, which is exactly the inspiration behind the Chinese name "Yue Zhi An Mian" (The Dark Side of the Moon).
All these details easily evoke DeepSeek-style associations: a young researcher, a non-internet-giant organization, approaching the frontier through architectural innovation, and breaking into global public opinion with a concentrated launch.
However, a similar debut does not mean identical commercial moves. What Kimi has been aiming for has always been clear: it is not trying to be the next "DeepSeek".
On March 25, Yang Zhilin delivered a speech titled "Open-Source AI: Accelerating the Exploration of the Upper Limit of Intelligence" at the plenary session of the 2026 Zhongguancun Forum Annual Conference, putting forward core judgments on the development of large models, and systematically disclosing Kimi's latest technical roadmap and industry value. During the speech, he summarized Moonshot AI's expansion directions into three points: improving token efficiency, extending context length, and organizing Agent clusters.
In less than four months, K3 has almost fully materialized these three judgments.
2
Kimi's goal: working along a continuous task thread
Token efficiency solves the cost problem, long context solves the capacity problem, and Agent clusters break through sequential execution bottlenecks. The directions Yang Zhilin planned for Moonshot AI all point to a single destination: Model competition is shifting from "a single answer" to "a continuously running task thread".
Previous relevant articles from HeFan Caijing, such as "Tencent, Alibaba, ByteDance: Rebuilding Office", "OpenAI Finds a New Job for Codex", and "The 'Forerunner' WPS: Fighting for Delivery Rights", all discussed this trend — AI competition needs to shift from "question-and-answer" to "task execution and delivery".
The end point of a chat is a reply, while the end point of work is a usable deliverable. This assistant, which once carried the functions of "consultation" and "search", now needs to go further and take on practical work.
Naturally, this will bring new requirements. For example, when completing a task, the assistant must remember the goal, read files and web pages, call terminals and browsers during the process, and save intermediate outputs. Compared with the past, this workflow becomes longer and more complex.
If the assistant makes a mistake at a certain step, it needs to know where to resume, instead of forcing the user to start all over again. If it forgets an important requirement or constraint while working, what might have been a minor experience flaw in the past could now mean a full restart of the entire task.
And a 1 million-token context does not equal a long thread. A longer context window is like a sufficiently large workbench, determining how much material you can spread out at one time; while a long thread is a continuously advancing project, which also depends on state preservation, tool feedback, task decomposition, and error recovery.
The brilliance of K3 is that, like DeepSeek, it uses technology to materialize its strategic goals.
For example, KDA first addresses the problem of rising costs for long tasks.
In October 2025, the Kimi Team published a related research paper titled "Kimi Linear: An Expressive, Efficient Attention Architecture". The paper already proposed a hybrid attention architecture combining KDA and MLA. On an experimental model with 48B total parameters and 3B activation parameters, the paper reported that KV Cache was reduced by up to 75%, and the decoding throughput at 1 million tokens reached about 6 times that of the full MLA baseline.
This is certainly not the actual measured result of K3, but it shows that if the cost of reading and saving historical information cannot be reduced, the longer the thread, the more difficult it is for the product to be viable.
And Attention Residuals solves the problem of information dilution in network depth.
Relevant papers show that the model can select representations from previous layers based on input, instead of adding them layer by layer in a fixed way. The verification also comes from the 48B model. We cannot claim that "the Agent will never forget the goal" based on this, but it does prove that this mechanism provides a more stable internal information flow for ultra-deep models.
And there is also Stable LatentMoE.
It allows each token to activate only 16 out of 896 experts, separating total capacity from single computation. Moonshot AI claims that these architectures and training recipes have increased scaling efficiency to about 2.5 times that of K2. As of press time, the technical report has not been released, so it is temporarily impossible to break down the contribution of each part.
However, making a single model more capable of holding information does not eliminate the limitations of sequential execution.
Kimi introduced Agent Swarm starting from K2.5, and by K2.6 in April 2026, it had expanded to a maximum of 300 sub-Agents and more than 4000 tool calls. The official stated that the speed of large-scale search is up to 4.5 times higher than that of a single Agent, and the current cluster is already driven by K3. Each sub-Agent retains its own context and only reports key conclusions to the coordinator. In other words, Moonshot AI is not only extending the thread, but also planning to split it into a temporarily organized team.
Products are also developing along the same direction.
Kimi Work, which launched in Beta on June 3, uses Kimi Code as its core, integrating local files, WebBridge, Skills, and permission control into the desktop. It can systematically decompose goals, call tools, and finally deliver documents, spreadsheets, presentations, or code.
K3 has also been deployed to Kimi.com, Kimi Work, Kimi Code, and APIs simultaneously. This means that Kimi is moving from a base model to a runtime and work entry point.
3
Here K3 and DeepSeek diverge
Entering the frontier game does not mean changing the rules of the game. What really made the market uneasy back then was that DeepSeek casually changed the price anchor point of the entire industry.
On May 6, 2024, DeepSeek released V2, triggering a domestic large model price war with extremely low API prices. Shortly after, in an exclusive interview with *Waves*, Liang Wenfeng said that the company only priced after calculating costs, following the principle of no subsidies and no pursuit of excessive profits.
This judgment turned into a complete set of actions on January 20, 2025.
On the day R1 and R1-Zero were released, DeepSeek launched 6 distilled versions ranging from 1.5B to 70B; the uncached input price of the API was $0.55, and the output price was $2.19 per million tokens. Capabilities, prices, weights, licenses, and models of different sizes were all available on the same day. Developers could not only call them, but also download, modify, deploy, and continue distilling them.
By the release of V4 Pro on April 24, 2026, this set of actions remained consistent.
V4 Pro has 1.6 trillion total parameters, 49B activation parameters, and a 1 million-token context. At launch, it simultaneously opened MIT-licensed weights, and is compatible with OpenAI Chat Completions and Anthropic interfaces. As of July 21, the official uncached input price is $0.435, and the output price is $0.87 per million tokens.
What DeepSeek reduced was not just the token price tag, but the threshold for acquiring, migrating, and modifying models.
On January 27, 2025, one week after R1's release, NVIDIA's market value evaporated by about $593 billion in one trading day, setting a record for the single-day market value loss of a US listed company at that time. Of course, this does not mean that computing resources are no longer important, but it fully records the market's fear: if near-frontier capabilities can be obtained at low cost, downloaded immediately, and compressed into models of different sizes, the price and capital expenditure narrative built around scarce intelligence must be recalculated.
The "DeepSeek Moment" represents that three things happen simultaneously: capabilities reach the frontier, prices plummet, and weights and distilled models are immediately available. Only then do powerful models shift from scarce products to compatible, modifiable supplies.
Kimi took a different path.
As of July 21, K3 has been deployed to Kimi Work, Kimi Code, and