Three consecutive strikes of China's open-source models, is Liang Wenfeng firing the last shot?
The full model weights of Kimi K3 are officially available on Hugging Face.
Just half an hour later, K3 received over 4,000 likes and topped the platform's trending list. Hugging Face co-founder and CEO Clem Delangue remarked that this was the fastest-growing release in the platform's history.
U.S. companies quickly followed suit. Guillermo Rauch, founder and CEO of Vercel, called K3 "the world's most capable open-weight model to date" and announced that Vercel had integrated K3 through a U.S.-based inference service provider. vLLM launched support on the very day the weights were made public, followed by platforms including Fireworks AI, Together AI, Modal, Baseten, and DigitalOcean.
Beneath this excitement lies a clear sense of urgency.
Just before K3's weights were released, OpenAI and Anthropic were emphasizing to Washington the distillation and security risks posed by China's open models. Jensen Huang publicly pushed back against this narrative, while companies including Microsoft, Meta, and NVIDIA jointly opposed premature restrictions on open-weight models.
The closer China's open-source models get to U.S. flagship models, the fiercer Silicon Valley's debate over open versus closed systems becomes.
Over the past three months, China's open-model ecosystem has delivered three major breakthroughs in succession.
The first came with the DeepSeek-V4 preview version, whose performance in some tests approaches that of U.S. closed-source flagship models at a far lower cost. The second was GLM-5.2, which supports a 1 million-token context window and focuses on complex, long-running Agent tasks. The third is Kimi K3, which just released its full weights: with a total parameter count of 2.8 trillion, it has entered the global first tier in multiple independent evaluations, outperforming Claude or GPT in some programming and long-task benchmarks.
These three breakthroughs directly target the three core advantages of U.S. closed-source models: pricing, long-task capabilities, and model scale.
Next, the unreleased official DeepSeek-V4 will deliver the fourth major push, aiming to challenge the last remaining stronghold of U.S. closed-source models.
A race has already begun: China's open models keep advancing, while U.S. closed-source companies are attempting to pre-emptively set up regulatory barriers.
Silicon Valley's "Divide"
On July 27, Moonshot AI released the full model weights for Kimi K3.
Almost simultaneously, Silicon Valley's debate over open models reached its peak intensity.
According to reports from The New York Times and Axios, OpenAI and Anthropic have recently warned U.S. policymakers that Chinese AI companies could distill lower-cost new models by extensively querying U.S. models and collecting their outputs. While APIs can restrict access by account and region, once model weights are made public, developers can download, modify, and deploy them independently, making it nearly impossible for the original company to reclaim control. The capabilities that U.S. closed-source companies spent billions to develop could quickly turn into infrastructure accessible to developers worldwide.
Jensen Huang publicly contradicted this stance. On July 22, he told Axios that strong Chinese open models "should be used" and that U.S. companies "absolutely should be allowed" to use Chinese models.
On July 24, 25 companies and organizations including NVIDIA, Microsoft, and Meta became the first signatories of the "Open Weights and U.S. AI Leadership" document, calling on Washington to avoid premature restrictions on open models.
OpenAI, Google, and SpaceX were not initially listed on the signatory list, but they joined the initiative over the following weekend.
This open letter opposes expanding the distillation controversy into blanket restrictions on open-weight models. As of July 27, among major U.S. frontier model companies, Anthropic had not yet signed the document.
Anthropic CEO Dario Amodei issued a dedicated response the same day, stating that his company had never advocated a total ban on open-weight models. He acknowledged that ordinary open models have public value, but argued that it is difficult to maintain safety guardrails when the most powerful models are made open. He therefore continues to support cracking down on industrial-scale distillation, restricting the flow of advanced chips to China, and requiring all high-capability models to undergo safety testing.
It is clear that OpenAI's position is equally nuanced: it supports the development of open models in the U.S., while continuing to warn Washington about potential unauthorized distillation practices by Chinese companies.
NVIDIA later partnered with Microsoft, IBM, Hugging Face, SpaceX AI, and other companies to establish the "Secure Open AI Alliance". The alliance's announcement noted that during Hugging Face's handling of a previous Agent boundary-violation security incident, some closed-source tools refused to assist in forensics due to security restrictions. The team ultimately ran GLM-5.2 locally, analyzed over 17,000 operations, and contained the intrusion. This led the alliance to conclude that defenders also need high-capability models that can be locally deployed, modified, and inspected.
The debate stems from fundamentally different business interests.
OpenAI and Anthropic monetize model capabilities through subscriptions and APIs: Claude Code alone reached an annualized revenue of over $2.5 billion as of February this year. NVIDIA sells chips, Microsoft sells cloud services, and Meta needs to expand its developer ecosystem — all of which hope to see more affordable, open models in the market.
Last January, teams behind Kimi, Qwen, and GLM appeared together at the Tsinghua AGI-Next Summit. Among China's "four leading open-source model developers", only DeepSeek was absent from the event. Their assessments were closely aligned: simply competing on basic chat capabilities can no longer create significant gaps, and the next round of competition will depend on who can enable Agents to handle longer and more complex tasks.
Six months later, DeepSeek, GLM, and Kimi have released successive updates that have forced Silicon Valley to publicly take sides. Kimi K3 is just the latest spark — the real driver of the controversy is the three major breakthroughs from China's open models over the past three months.
China's "Three Major Breakthroughs"
In 2026, China's open models have caught up in three consecutive waves, each dismantling one of the exclusive advantages of U.S. closed-source models.
The "first breakthrough" came from DeepSeek.
On April 24, DeepSeek released the V4 preview version, with model weights made available simultaneously. The V4-Pro has a total of 1.6 trillion parameters, activates 49 billion parameters per inference call, and supports a 1 million-token context window.
In self-test results published on DeepSeek's model card, V4-Pro achieved a SWE-bench Verified score of 80.6% at maximum inference intensity, nearly matching Claude Opus 4.6's 80.8%. Its BrowseComp score reached 83.4%, close to Claude Opus 4.6's 83.7% and higher than GPT-5.4's 82.7%.
The most disruptive factor is pricing. On May 23, DeepSeek announced that it would make its previously limited-time 75% discount a permanent price, effective after May 31. Currently, it charges $0.435 per million tokens for non-cached input and $0.87 per million tokens for output.
This means developers can now access an open model with near-flagship performance at a fraction of the cost of U.S. closed-source alternatives.
The "second breakthrough" came from Zhipu AI.
On June 16, Zhipu AI released GLM-5.2, which also supports a 1 million-token context window and is open-weight under the MIT license.
Its focus has shifted from answering questions to sustained execution of long-running tasks: reading large codebases, iteratively calling tools, and completing software engineering workflows across hundreds of operational rounds.
According to aggregated best-in-class test results from Zhipu AI, GLM-5.2 reached 82.7% in Terminal-Bench 2.1, slightly lower than GPT-5.5's 83.4% but higher than Claude Opus 4.8's 78.9%.
Even accounting for differences in Agent tooling across models, these results demonstrate that open models are now capable of handling real terminal operations and complex workflows, eroding the lead closed-source models once held in Agent capabilities.
The third breakthrough came from Kimi K3.
On July 27, Moonshot AI released K3's full weights. According to the company, this is the world's first open-weight model with a 3-trillion-parameter-scale architecture: it has 2.8 trillion total parameters, activates 104 billion parameters per call, and supports images, videos, and a 1 million-token context window. K3 combines parameter scale, multimodality, and long-task capabilities that were rarely seen together in previous open models.
In test results published by Moonshot AI, K3 scored 88.3% in Terminal-Bench 2.1, surpassing Claude Fable 5's 88.0% and just 0.5 percentage points behind GPT-5.6 Sol's 88.8%.
In SWE-Marathon, a benchmark designed to test ultra-long software engineering tasks, K3 scored 42 points, outperforming both GPT-5.6 Sol (39 points) and Claude Opus 4.8 (40 points).
With these three breakthroughs, pricing, long-task capabilities, and model scale are no longer exclusive advantages of closed-source models.
However, most of these results come from vendor-conducted tests, and some comparisons mix different Agent tools and runtime environments. They have not yet proven that China's open models can match the overall stability of U.S. closed-source flagship models in real-world projects and sustained, repeated operations.
This is the final barrier that the official DeepSeek-V4 is designed to break through.
What Remaining Barrier Does V4 Aim to Overcome?
Since some evaluation metrics have already caught up, why can U.S. closed-source models still charge ten or even dozens of times more for API access?
Moonshot AI provided a candid answer in the official K3 model card. Even though K3 has achieved top-tier performance in multiple benchmarks, there is still a "noticeable gap" in overall user experience compared to Claude Fable 5 and GPT-5.6 Sol.
GLM-5.2 faces similar issues. It reached a maximum score of 82.7% in Terminal-Bench 2.1, approaching top U.S. model performance. But in SWE-Marathon, a benchmark dedicated to ultra-long software engineering tasks, Zhipu AI's published results show GLM-5.2 only scores 13 points, while Claude Opus 4.8 scores 26 points.
For enterprises, these small errors directly translate into tangible costs.
When companies select models, they must account for both API fees and the time spent reworking failed outputs.
DeepSeek-V4-Pro charges $0.435 per million tokens for non-cached input and $0.87 for output, while Claude Opus 5 charges $5 and $25 respectively — its output price is nearly 29 times that of DeepSeek. This is a massive price advantage, but if a long-running task fails at the final stage, engineers will have to re-examine, modify, and re-run the entire workflow, quickly erasing any savings from lower API costs.
Therefore, the official V4 release needs to deliver not only higher benchmark scores, but also stable tool calling, long-context processing, and multi-turn modification capabilities that hold up under repeated use. It aims to give developers the confidence to entrust large codebases, complex research, and long-duration Agent tasks to open models.
As of now, DeepSeek has not announced the official V4 release date or full specifications. In a recent circulated investor discussion, DeepSeek founder Liang Wenfeng stated that the company will most likely continue to open its most powerful models, with commercialization focused only on "reasonable profits".
If this commitment holds, the official V4 will continue DeepSeek's most potent formula: performance comparable to U.S. flagship models, far lower pricing, and open weights.
With this context, the recent moves by OpenAI and Anthropic become far easier to understand.
Three major breakthroughs have already been delivered, and these companies do not want to wait until the official V4 launch to respond. That is why, immediately after Kimi K3 went open-source, the two companies pushed Washington to launch investigations into model distillation, and discuss potential sanctions and usage restrictions. Once model weights are widely downloaded and deployed, any subsequent restrictions will be far less effective.
It is now a race to see which arrives first: the official DeepSeek-V4, or new restrictions from Washington.
The full impact of this fourth breakthrough will depend half on whether DeepSeek can deliver a stable, low-cost, fully open official version, and half on whether the faction advocating restrictions on China's open models succeeds in advancing its agenda.
This article is from the WeChat public account "