Wall Street lifts Kimi only to drop it on its own feet
48 hours.
A Chinese company releases its model.
Global AI stocks plummet collectively. Nvidia falls, semiconductor stocks drop... Trillions of dollars in global market value evaporate.
A Goldman Sachs partner directly stated on a client call: The era of computing power expansion may be over.
A memo sent by JPMorgan Chase to clients refers to this event as DeepSeek 2.0.
Everyone is asking the same question: Is Kimi K3 really that powerful?
To be honest, I don't think so.
Kimi is a good model. But it is not enough to shake the global capital market. What has truly been shaken is not Nvidia, but the narrative that Wall Street has believed in for the past three years.
Kimi is just the hand that topples the dominoes.
I. What Exactly Is Wall Street Afraid Of
To understand this panic, we must first clarify what Wall Street truly believed in over the past three years.
In essence, they subscribed to a fundamental law of large language models.
The more powerful the model, the larger its parameters, the more expensive its training, the more GPUs it requires, the higher the capital expenditure, and the more valuable Nvidia becomes.
This has also been the largest financial narrative of the entire AI era.
After ChatGPT exploded in popularity in 2023, this formula persisted all the way to mid-2026. Nvidia's market value surged from 1 trillion US dollars to 5 trillion US dollars. Microsoft, Meta, Alphabet, and Amazon repeated the same line at every earnings call: we will continue to increase investment in AI infrastructure.
The valuation logic was built on only one assumption: doing AI requires burning money. The more you burn, the deeper your moat.
For three years, this formula was never falsified even once.
Until DeepSeek emerged.
In January 2025, DeepSeek-R1 achieved inference performance close to GPT-4 at a cost far lower than expected. On that day, Nvidia lost 589 billion US dollars in market value, setting a historical record for the US stock market.
And what happened next?
Six months later, Microsoft continued its investments. Meta increased its budget by 20%. CoreWeave's stock price doubled after its IPO. Everyone kept placing their bets.
Wall Street breathed a sigh of relief, and DeepSeek was dismissed as a false alarm.
Then Kimi K3 arrived.
The data this time is even more glaring. It topped the LMArena front-end code leaderboard with a score of 1679, surpassing Claude Fable 5. Its overall ranking jumped directly from 18th place in the previous generation to 1st. It took 6 first places in 7 front-end sub-categories. The open-weight release on July 27 reportedly cost only 60% of that of top-tier US models.
Goldman Sachs posed this rhetorical question in its internal memo:
How can a Chinese lab that cannot access large-scale computing power narrow the gap so rapidly through architectural innovation, synthetic data, reinforcement learning, and post-training?
The moment this question was asked, Wall Street's belief began to crumble.
Previously, everyone took it for granted that Scale meant GPUs. Today, Scale has become the combined force of GPUs, algorithms, data, post-training, and reinforcement learning.
Computing power is no longer the only answer, and that is what Wall Street is truly afraid of.
What Wall Street fears is not Kimi, but Nvidia. More accurately, it fears the pricing logic that AI capital expenditure will keep expanding forever.
Silicon Valley giants are projected to spend 700 billion US dollars on capital expenditure in 2026. TSMC just announced a 64 billion US dollar capital expenditure for 2026, exceeding market expectations. This money is supposed to be converted into chips, data centers, and orders for Nvidia. The premise is that large language models must keep growing larger and keep burning money.
After Kimi K3 was released, people realized that this is not necessarily the case.
If more and more companies can build models of the same level with fewer GPUs and lower costs, will the global GPU procurement volume still need to be that large? Will the return on TSMC's 64 billion US dollar investment be discounted? Should Nvidia's forward revenue projections be revised down?
Now, all the valuation models for computing power expansion built over the past three years have to be recalculated.
Kimi just made everyone truly start asking a question for the first time: will GPUs still be sold in such large volumes?
As long as Wall Street begins to ponder this question, stock prices are bound to fluctuate.
II. Why Is It China Again?
There is a deeper layer that many people dare not state openly.
Why is it China again?
DeepSeek made Wall Street realize for the first time that China can catch up. Kimi told Wall Street the second time that China is still catching up and has not stopped.
The first occurrence can be called accidental, attributed to a few geniuses or pure luck. When low-cost overtaking happens twice at different stages and through different methods, capital has to recalculate the probabilities.
Once this probability is revised, everything changes.
US giants believe in achieving miracles through sheer brute force. With abundant funds and numerous GPUs, they can cover up engineering flaws by throwing even more computing power at the problem. This approach is actually inefficient, but no one cared about it in a resource-abundant environment.
Leading Chinese companies have faced the exact opposite environment from day one. They cannot obtain clusters of 100,000 GPUs, and their financing is not that generous, so they can only maximize the utilization of the computing power they have.
This forced them to develop two core capabilities.
The first is synthetic data. Using the minimum number of high-quality samples to push the upper limit of logic to the highest level.
The second is post-training combined with reinforcement learning. Using engineering fine-tuning and strategy optimization to compensate for deficiencies in underlying pre-training.
These two capabilities are precisely the areas that the computing power faith has long neglected.
The result is that when the pursuer can achieve top-tier performance at 40% or even lower cost, the 100,000-GPU barrier built by US giants becomes a cost burden instead.
This is Wall Street's real systemic fear. It is not that a Chinese company has built a good model, but that this low-cost path has been proven feasible twice. The third and fourth instances may already be on the way.
III. But Wall Street Might Have Misjudged Again
Just as Wall Street was selling off stocks and discussing whether the computing power era is coming to an end, reality delivered a reversal.
Last night, Kimi's official announcement stated that after the release of K3, the number of requests far exceeded expectations, and the cluster is approaching its carrying limit. New user subscriptions for the C-end will be suspended immediately, and full-speed capacity expansion is underway. To match the computing power supply, the most power-hungry Kimi Code feature will be separated from the main subscription benefits and sold separately.
This news dealt Wall Street a heavy slap in the face.
It has brought a repeatedly verified law of the technology industry to the forefront. Cost reduction never reduces demand; instead, it causes demand to explode.
Examples of this are everywhere.
When computers became cheaper, the number of computers doubled. When internet access became cheaper, the number of websites exploded. When cloud computing became cheaper, the total number of servers worldwide became far larger than before.
What will happen now that models are becoming cheaper?
More Agents will be created, the volume of AI application calls will surge, inference demand will skyrocket, and even Kimi itself cannot handle the load.
This phenomenon is known in economics as the Jevons Paradox. A 19th-century British economist discovered that after the efficiency of steam engines improved, the total national coal consumption in the UK actually rose. Higher efficiency lowered the usage threshold, leading to widespread adoption of steam engines across numerous industries, and the overall demand expansion offset the energy savings per individual device.
Today's GPUs are the coal of the 19th century. GPU demand will not disappear; what changes is its role.
In the past, the main buyers of GPUs were training teams. Tens of thousands of GPUs were concentrated in closed clusters for months of training, and many of them were intermittently idle after training was completed.
In the future, the main buyers will be inference services. 24/7 Agent services, local models on every mobile phone, every vehicle, and every robot, video generation, and continuous computation for million-token windows.
The pattern of computing power consumption is shifting from low-frequency, massive, and concentrated to high-frequency, massive, and ubiquitous.
The total volume may not decrease; it may even become larger. But the order of who benefits and how they benefit will be completely reshuffled.
Goldman Sachs claims that the era of computing power expansion is nearing its end. At least within a two-year timeframe, this judgment is most likely wrong.
A more accurate statement should be that the logic of computing power expansion has changed. The dominance has shifted from training to inference.
Nvidia will not collapse as a result. But the days of blindly buying Nvidia stocks as a sure bet, which have lasted for the past three years, may truly be over.
IV. Kimi Is Just the Mirror
Looking back at this sell-off, the role of Kimi K3 is very clear.
It is the trigger, not the cause.
You will understand this by looking at a few key data points.
The Philadelphia Semiconductor Index has risen by 68% this year. After such a large increase, a deep correction was inevitable. The S&P 500 Equal Weight Index hit a historical high on the same day. The money did not disappear; it just moved from tech stocks to other sectors.
Data from EPFR shows that US corporate insiders sold 77.6 billion US dollars worth of stocks in the first half of the year, the second-highest figure in 20 years. Smart money has long been retreating.
TSMC's Q2 net profit rose 77% year-on-year, and its revenue forecast was raised. However, the 64 billion US dollar capital expenditure that exceeded expectations instead made the market anxious.
Alphabet's Gemini 3.5 Pro was delayed for several months, and the implementation pace of its AI products failed to meet expectations.
All these factors combined have made the market extremely tense, just waiting for a reason to release the pressure.
Kimi K3 just provided that reason, giving Wall Street a brick to hit its own foot.
Strictly speaking, this sell-off is more like the loosening of a crowded trade, rather than the sudden collapse of the AI narrative. The funds that have been concentrated in Nvidia over the past three years are beginning to flow elsewhere.
Kimi did not change AI; Kimi changed how capital understands AI.
V. Old Maps, New Continents
From today's perspective, the competitive logic of the AI industry is rapidly shifting gears.
The first phase ran from 2023 to 2025, with the core focus on computing power reserves. The party with more GPUs and the ability to train larger models was the winner. The winners in this phase were Nvidia, TSMC, and a few giants that could afford tens of billions of dollars in training budgets.
The second phase starts at the end of 2025, with the core shifting to efficiency. Whoever can build models of the same level with fewer resources, turn models into products faster, and establish an advantage in inference costs will seize the initiative in the next round.
The problem with Wall Street has never been that it is not smart, but that it is too fond of using a single simple narrative to explain a complex era.
Three years ago, AI was framed as a GPU story. Valuations revolved around GPUs, and investment positions were concentrated around GPUs.
Today, this narrative has for the first time shown obvious logical blind spots.
A blind spot does not mean the end; it means the narrative needs to be upgraded.
The focus shifts from who has the most GPUs to who can use GPUs the best. The transition is from hardware dividends to engineering capabilities and system efficiency, and from capital-intensive development to technology-intensive and application-intensive development.
DeepSeek made Wall Street start to doubt. Kimi made Wall Street start to correct its views.
A message beyond the layout:
If one day in the future we look back, Kimi K3 may not necessarily become the most important model. But it will very likely become a landmark.
Not because it surpassed anyone, but because it made global capital realize for the first time that AI competition is no longer about who owns more resources. Now it is about who can turn the same resources into greater capabilities.
These are two completely different industrial eras.
In the previous era, money determined speed. In the next era, efficiency determines value.
This article is from the WeChat official account "Beyond the Layout", author: Huahua, published with authorization from 36Kr.