Kimi triggers the circuit breaker, Yang Zhilin also reaches for new heights
K3 became a viral hit, but Kimi had no choice but to hit the pause button.
Late at night on July 19, the Moonshot AI Kimi team released an announcement stating that within 48 hours of Kimi K3's launch, user requests had far exceeded projections. Due to tight computing power resources, new user subscriptions for Kimi would be temporarily suspended, and limited computing power would be prioritized to serve existing subscribed users.
The membership system was subsequently split into Kimi's core benefits and Code-specific benefits to allocate resources precisely. According to the official response released shortly afterward, Kimi did not adjust the membership system; it merely rebalanced the user experience under the constraints of limited computing power.
In other words, K3 became so popular that it triggered a computing power "circuit breaker", forcing the platform to politely decline new users and focus on serving existing ones.
Kimi's "Tao's Law"
Currently, all major model manufacturers are striving to expand their user base, attracting new users through cost-effective measures such as free access and price reductions. In comparison, Kimi's counterintuitive move seems somewhat "show-off".
However, the stronger the model's capabilities, the longer the context window, and the more frequently users make calls, the higher the consumption of inference computing power. When "K3's request volume is approaching the load limit of the existing computing power cluster", Kimi has encountered the same "Doubao-style" dilemma: the more C-end users there are, the less obvious their contribution to revenue becomes, yet the GPUs cannot keep up.
As Moonshot AI's first 3-trillion-parameter MoE model, K3 has a total parameter scale of 2.8 trillion. Conventional understanding suggests that a model of this size would consume computing power exponentially with each output. However, supported by its unique core architecture, it achieves the effect of "Tao's Law" proposed by Huawei in the semiconductor field.
This architecture is also regarded as a super engine that implements three key technologies: KDA (Kimi Delta Attention) mixed linear attention, Attention Residuals, and high-sparsity MoE, pushing the efficiency of converting every unit of computing power into model performance to the extreme.
For a long time, the semiconductor industry has relied on Moore's Law to improve performance by shrinking transistor sizes. However, as advanced manufacturing processes approach physical limits and costs skyrocket, the growth model of simply stacking hardware has hit a bottleneck.
Huawei proposed "Tao's Law" this year, breaking away from the traditional idea of "spatial miniaturization" and turning to "temporal miniaturization": through logical folding, three-dimensional stacking, and full-link architecture optimization, it compresses signal transmission latency and reduces redundant data movement, achieving an energy efficiency leap on mature manufacturing processes that rivals that of advanced processes.
Its core methodology is: under hardware or cost constraints, achieve performance leaps through system-level architecture innovation.
In my opinion, the KDA mixed linear attention mechanism adopted by K3 is precisely the practice of "Tao's Law" in the AI field.
It is well known that the Transformer architecture consumes the most computing power in its attention mechanism. Standard attention grows quadratically with the sequence length, and in million-level context tasks, most computing power is consumed maintaining the attention matrix of long sequences. The KDA mechanism reduces the computational complexity of most attention layers from quadratic to linear.
According to the previous technical report released by Moonshot AI, by adopting a mixed ratio of KDA and MLA on the experimental model, KV cache usage is reduced by up to 75%, and the decoding throughput under a 1 million context window is increased to 6 times the original. Combined with Attention Residuals and high-sparsity MoE technologies, training efficiency is improved by approximately 25%, with additional costs of less than 2%.
This is a reconstruction of "model production efficiency". If the mainstream Silicon Valley Transformer route is a "muscle-over-mind" fuel vehicle, KDA is more like a hybrid engine that pursues "ultimate thermal efficiency".
A report from SemiAnalysis points out that KDA increases inference speed by 300% while maintaining performance close to Transformer in long-context processing. This means that with the same computing power input, K3 can produce more effective intelligence.
Actual test feedback from the developer community shows that K3 performs outstandingly in long-process Agent tasks and code generation, possessing the strength to compete with the world's leading models.
Fundamentally, K3 still follows the Scaling Law, but it targets the waste of crude algorithms. While Silicon Valley AI companies are still hoarding GPUs and stacking computing power, Kimi reconstructs the algorithm link and streamlines computational redundancy to achieve a balance between performance and efficiency under limited computing power, embarking on a path that maximizes "intelligence output per watt".
This surprised Wall Street observers and investors. Just as they were astonished by DeepSeek, they once again questioned whether the hundreds of billions of dollars invested in AI in Silicon Valley would yield corresponding returns and whether the United States' leading position in AI was being eroded.
At the same time, enterprise users and developers are starting to pay attention to AI usage costs. Some developers have already adopted "Model Routing" services, which automatically select the most cost-effective and efficient AI model for different tasks, including of course Chinese models such as Kimi.
Yang Zhilin's High-Reaching Ambition
In retrospect, Kimi's computing power being overwhelmed is a side effect of technological success: the improvement of model efficiency reduces the cost per use, which in turn stimulates a sharp surge in total demand.
It is worth mentioning that on July 11, Tang Jie, founder of Zhipu AI, announced the "Touch High" initiative, stating bluntly that "failing to reach the top means failure". He also made it clear that the company will not pursue short-term application monetization in the next two years, but concentrate resources on tackling the four core directions of AGI, and plans to raise funds to invest tens of billions of resources to break through mechanical interpretability technology.
If Zhipu AI's "high-reaching" move is a do-or-die battle after its listing, it hopes to consolidate the narrative of "the world's first large model public company" by challenging the technological ceiling and alleviate real anxieties.
I also see Kimi's "high-reaching" efforts in three dimensions from Yang Zhilin's decisions:
First, reaching high in computing power: shifting from a traffic-centric mindset to a value-centric mindset. Public information shows that Kimi's API business has accounted for over 70% of its revenue, which is vastly different from its previous C-end subscription model. Prioritizing existing users and turning away new ones essentially tilts limited computing power toward high-value scenarios, preventing general C-end traffic from crowding out the quota of the B-end business that is already the revenue pillar.
The downsides are also obvious. For example, the execution criteria for "prioritizing existing users" are not transparent. Which requests receive priority response and which are downgraded—any mishandling will erode user trust.
In the long run, for Kimi to complete its transformation from a tech star to a mature leading AI enterprise, it must consolidate infrastructure such as elastic computing power pools, hierarchical response mechanisms, and dynamic scheduling capabilities.
Second, reaching high in pricing: Kimi is conducting a stress test shifting from cheap all-you-can-eat to value-based tiering. The "split of Kimi's core benefits and Code-specific benefits" mentioned in the announcement is not only a short-term measure to cope with computing power shortages, but may also herald the beginning of the reconstruction of its pricing system.
In the past, whether users were coding, researching, or chatting casually, they paid the same fee and consumed the same amount of computing power. As the proportion of Code users with high computing power consumption rises, this model will inevitably lead to structural mismatches.
Kimi takes this opportunity to split its membership benefits, re-pricing based on usage scenarios: light users enjoy basic services, while heavy users pay a premium for high-value capabilities.
This is an attempt to seize pricing power for model products by segmenting value users. Currently, K3's pricing has reached $2.3 per million Tokens. Although it is lower than that of top overseas models, it has set a new record for domestic models.
Perhaps Kimi also wants to bid farewell to the low-cost model and convey a message to the market: high-performance models have pricing power. Next, the competition among large models will shift from "competing on parameters" to "competing on infrastructure" and "competing on cost-based pricing".
This step is more difficult than topping the rankings, and it is also closer to the essence of business.
Third, reaching high in status: Kimi's identity transition from a geopolitical variable to a pricing anchor.
K3 quickly attracted global attention after its release. In the latest "Intelligence Index" released by the independent AI model evaluation agency Artificial Analysis, its composite score reached 57 points, ranking third in the world, second only to Claude Fable 5 (60 points) and GPT-5.6 Sol (59 points), and ahead of Claude Opus 4.8 (56 points).
This achievement also breaks the industry consensus that "open-source models lag behind closed-source models by half a generation".
At the same time, K3's influence has spilled over beyond the tech circle. US tech investment institutions regard it as an "inflection point" in AI development, believing that high-quality open-source models are accelerating capability diffusion. A professor of computer science at the University of California, Berkeley claimed that K3 has narrowed the gap between China's open-source models and the US's advanced models to 2 to 3 months.
The capital market's reaction was equally intense. On the first trading day after K3's release, Nvidia's market value evaporated by approximately $111.1 billion in a single day; JPMorgan Chase strategists directly referred to it as "DeepSeek 2.0" in their report.
This acceleration is what changes the pricing logic.
In addition, Moonshot AI's ARR (Annual Recurring Revenue) has approached $300 million by June, with overseas growth rate hitting 400%. After a 60% price increase, revenue has instead risen. The combination of these three sets of data indicates that Kimi's "open-source + efficiency" model can achieve large-scale monetization.
Coupled with disclosures that Moonshot AI is seeking to list in Hong Kong in as fast as six months, its valuation has soared from $4.3 billion to $31.5 billion in half a year, a surge of over 7 times.
All these are the capital market's "recognition" of a non-mainstream Silicon Valley route. When they bet on "SOTA that can make the unit economic model work", it also shows that Kimi has its own independent valuation logic, no longer needing to benchmark against OpenAI or Anthropic to prove its value.
Moonshot AI Source: Online materials
However, K3 still has a long way to go before reaching AGI. For example, in benchmark testing, although K3 is only 3 points lower than Fable 5, there is still a noticeable gap behind the scores.
In addition, Kimi also acknowledged two deployment-level flaws in its technical report: First, K3 retains the complete thinking history during training. If the caller does not correctly send back the previous reasoning process, or switches to K3 from another model in the middle of a session, the output quality will drop significantly;
Second, K3 will be "overly proactive" when the task is ambiguous, tending to make decisions for users on its own. This kind of "arbitrary action" can easily trigger cascading errors in engineering processes that require precise execution.
There is another easily overlooked but extremely important issue: hallucinations. Artificial Analysis specifically pointed out in its evaluation that K3's hallucination rate has actually increased compared to the previous generation K2.6.
To some extent, K3's glamour and hidden worries are a microcosm of the industry-wide imbalance between computing power supply and demand, as well as the collective pain of China's AI industry under the heavy pressures of computing power bottlenecks, unclosed commercialization loops, and excessive user expectations.
Every overload and every "circuit breaker" of these AI manufacturers is a hurdle that China's AI "high-reaching" efforts need to overcome. What is certain is that a new competition cycle has begun, and large models will shift from competing on capabilities to competing on computing power.
Whoever can build advantages in infrastructure such as algorithms and computing power reserves will have the last laugh and define the survival standards for the next generation of AI companies.
References:
Letterboard, "48 Hours After K3's Release"
Jinduan, "It's All Kimi's Fault"
This article is from the WeChat public account "Tang Chen Classmate", author: Tang Chen, published with authorization from 36Kr.