HomeArticle

The collective stealthy overseas expansion of Chinese large models

王智远2026-09-04 12:40
Silicon Valley has downloaded Chinese models to local hard drives.

On June 12, 2026, the U.S. Department of Commerce issued an export control order to Anthropic.

The order requires that starting from that day, no foreign nationals are allowed to access Fable 5 and Mythos 5, the two most powerful models, and even Anthropic's own non-U.S. overseas employees are not permitted to do so.

Unable to verify the nationality of each user within seconds, Anthropic simply made a more thorough decision: to take the service offline globally, making it unavailable to anyone.

I guess most of you have seen this piece of information some time earlier, right?

On that very day, the world's most powerful model became inaccessible to users across the globe. Enterprise customers found that the core service vanished overnight without any prior warning or appeal channel, including service providers in the finance, healthcare and software sectors.

18 days later, the control was lifted and the model was back online with attached conditions: Anthropic promised to cooperate with the U.S. government's review of new model releases, voluntarily report security risks, and the Department of Commerce reserved the right to pull the plug again at any time.

When U.S. media reviewed this incident, they used the term "regulatory kill-switch". In just one afternoon, it turned from a theoretical concept into an operational reality.

At that time, a group of cybersecurity researchers jointly issued a public statement opposing the order, pointing out that the one-size-fits-all restriction would first harm local cybersecurity defense teams. There was also a more pragmatic calculation: the days when the service was shut down handed over a valuable time window to open-source peers on the other side of the ocean.

Interestingly, other models under Anthropic were not affected at all throughout the process, and only the two most powerful ones were shut down for 18 days. When the switch is pulled, it is always the most powerful targets that get chosen.

The reason I am talking about this incident is to lead to a bigger change that follows.

In the past six months, several of the world's most valuable AI companies, including Cursor, Harvey and Thomson Reuters, have successively replaced the underlying foundation of their self-developed models from U.S. closed-source models to Chinese open-source models.

Among them, Harvey is a company invested by OpenAI, with a valuation of 11 billion U.S. dollars and serving more than 2,000 institutions.

To understand this collective choice, we first need to figure out what they are afraid of.

They are afraid of the kill switches, and there are two of them in total. The first one is political, which the U.S. government has just demonstrated how to operate, and the second one is commercial, held by API suppliers.

The second switch stays quiet most of the time, but once triggered, it will lead to price adjustments, clause revisions and changes of product direction.

If you run a company that relies on calling third-party APIs to survive, you have no choice but to accept the supplier's price increase today, wait for their service adjustments tomorrow, and you will not even have a place to appeal when they enter your business field to compete with you the day after tomorrow.

Both OpenAI and Anthropic are expanding their business to vertical industries such as legal services and programming.

As a legal AI company with investment from OpenAI, Harvey understands better than anyone else that the supplier whose API it calls every day will sooner or later come to seize its own market.

Even the suppliers themselves cannot avoid this political kill switch. When OpenAI's new flagship model was just released, its top-tier cybersecurity capabilities were not fully opened to the public, but only provided to a group of pre-vetted institutions. The most advanced intelligence comes with locks from the very beginning of its birth.

The technical leader of Thomson Reuters expressed this concern more directly. He said:

In the past, calling other people's models was like renting a house: you have a roof over your head, someone else is responsible for maintenance, and it feels comfortable to live in. But the house is not yours, no matter how long you live in it, you will never accumulate your own assets.

Developing a model on your own is like buying a house: the upfront investment is large, you have to calculate all the costs yourself, but the house will belong to you completely from then on. The person who said this is the technical lead of the parent company of Reuters, a group that has been in the information business for more than a hundred years.

The kill switch cannot be pulled several times a year, but the "rent" has to be paid every month. The intelligence you rent will always have its lifeline controlled by others. Where else can they go?

......

They are turning to models that can be downloaded to local hard drives. Almost all the most competitive models of this type come from China.

The first group to take action was developers. In March, Cursor released its self-developed model Composer 2, whose official blog was full of descriptions about continuous pre-training and reinforcement learning, without a single word mentioning its underlying foundation.

At that time, the company had a valuation of around 30 billion U.S. dollars and was the most popular programming tool in Silicon Valley.

It encountered an accident within 24 hours after the release: some developers checked the API logs and found a line of model ID: kimi-k2p5-rl-0317-s515. The prefix corresponds to Kimi K2.5, "rl" refers to reinforcement learning, and "0317" is the training date.

Elon Musk reposted the content personally, confirming that it was Kimi K2.5.

The team from Moonshot AI added further proof that the tokenizer was completely identical, and they had never received any authorization application for it.

They had valid reasons to be angry: the license agreement of K2.5 clearly states that products with monthly active users exceeding 100 million or monthly revenue exceeding 20 million U.S. dollars must prominently mark Kimi K2.5 on the interface.

At that time, Cursor had a monthly revenue of about 167 million U.S. dollars, 8 times the threshold, but its interface had no related marks at all, only showing the name Composer 2.

Finally, the founder of Cursor admitted that after evaluating all base models, they found Kimi K2.5 was the most powerful, and the failure to mark it was a mistake.

The dispute lasted for three days before the official Kimi team announced that it was an authorized cooperation, and Cursor accessed the model through the Fireworks platform with complete formalities, bringing the incident to a decent end.

The Cursor incident in March was not an isolated case. Two days before it, Japanese e-commerce giant Rakuten released its so-called most powerful domestic large model, but less than 12 hours after launch, the community found out that its base model was DeepSeek-V3, and it did not even keep the original license file until it was exposed by the public.

The word "self-developed" has suddenly become the most expensive ornament in the industry.

Two months later, when it released Composer 2.5, it wrote the base model information clearly in the first sentence of the announcement, and never tried to hide it again. The same company that concealed the details in March took the initiative to make a full disclosure in May, and the industry trend changed even faster than the iteration speed of models.

The story does not end here.

In June, Elon Musk's SpaceX officially announced the acquisition of Cursor for 60 billion U.S. dollars, with the delivery to be completed in August. The person who personally ridiculed the incident in March finally brought the company that uses Kimi under his own banner.

After the developer community, the legal industry followed suit. In August, Harvey officially released its self-developed model, and even wrote the base model information in the title without hiding it. The release note clearly states that the visible part of the consideration is cost control, and the deeper core reason is autonomous control.

Using 150 NVIDIA B300 GPUs for two months of training, it increased the full pass rate of Kimi K3 on legal tasks from 10.8% to 19.7%, doubling the number of tasks that can be completed fully, while the cost per task remained the same as that of the base model.

At the same time, OpenAI used more than 100,000 GPUs to train its new flagship model. An application company can train a professional model with just over a hundred GPUs, which means the threshold of replacing the base model is far lower than most people imagined.

The old-money enterprise came last and kept a low profile. At the end of August, Thomson Reuters invested 40 million U.S. dollars to develop its own large model.

The company has accumulated 150 years of judicial precedents and regulations, and its Westlaw database is more valuable than the total data volume of many other industries. Its official press release was written very carefully, only mentioning that the model is based on a "powerful open-source foundation", without specifying which one it is.

When its CTO was asked about the base model in an interview, he admitted that it was Alibaba's Qwen3.5.

Before using it, they worked with Imperial College London for several months of recalibration to convert the third-party base model into their own intermediate model, and named it Snowdon.

They spent the money and finished the work, but refused to mention the source of the base model. This is the real price of sensitivity over intellectual property attribution.

At this point, some people may ask: What about the open-source models developed in the U.S.? Why not download those instead?

Three years ago, the open-source model world was dominated by Llama. After Meta open-sourced it, developers all over the world focused on it. Now the leading position has changed hands, and the top names on the ranking list have shifted from San Francisco to Hangzhou and Beijing.

The U.S. did not fail to notice this trend.

Meta once closed its source code in the first half of the year, only selling access to its flagship model without opening the weights, but it changed its mind on August 10. Mark Zuckerberg published a long article to embrace open source again, releasing a 30B local model and promising to gradually open the weights of its flagship model.

OpenAI took earlier action. Last August, it released its first open-source model since GPT-2, with 120 billion parameters.

However, the top of the open-source ranking list is all occupied by Chinese teams.

European companies calculate more carefully. A partner at a German consulting firm said that running Chinese open-source weights on local European servers keeps data in their own hands, which is more in line with the concept of data sovereignty than calling U.S. APIs that can be remotely adjusted in price or revoked at any time.

Siemens has integrated Alibaba's model into its industrial software. They are not necessarily pro-China, they just do not want to bet on any third party's kill switch anymore.

On the model aggregation platform OpenRouter, the proportion of traffic from U.S. companies accessing Chinese models has been rising all the way this year. It has exceeded 30% every week since February, and the peak once reached 46%.

The trend of price difference is even more subtle. Before August, Chinese models were known for being extremely cheap. For the same task, their price was only a fraction of that of U.S. flagship models. No matter how wealthy Silicon Valley enterprises are, they cannot afford such a huge cost gap accumulated on the monthly bill.

On August 17, the trend changed. DeepSeek raised the output price of its flagship model by 350% at one go, charging 27 yuan per million tokens during peak hours.

In the same week, OpenAI cut the price of its flagship model by 80%, almost giving it away for free. Chinese model providers raised prices while U.S. providers cut prices. The price gap is narrowing, but the download volume of Chinese models has not declined.

The annualized revenue of Kimi has nearly tripled in three months. The number of users downloading its models is rising, and the revenue side is also doing very well.

Such a large-scale migration of users cannot be hidden from both sides.

......

Washington took the first action. At the end of July, Republican Senator Tom Cotton wrote a long letter to the U.S. Secretary of Commerce.

The letter cited a report saying that nearly 80% of U.S. startups build their applications on Chinese models.

He specifically mentioned that Airbnb's customer service robot uses Alibaba's Qwen, and Cursor's programming base model is Kimi. He also added that defense contractors are using Cursor to write software for the U.S. Department of War, and there may be backdoors embedded in the model weights.

He suggested expanding the ban from DeepSeek to all Chinese models, and prohibiting all government contractors from using any of them.

Tom Cotton's remarks were very radical, but this letter has remained a proposal so far and has not been implemented.

AT&T, the largest telecom company in the U.S., runs 40% of its internal AI queries on NVIDIA's open-source models, but it publicly stated that it has not used any Chinese models, and is still evaluating the risks of DeepSeek and Kimi.

While enjoying the benefits of open source, they are avoiding the most competitive batch of open-source models. The phrase "want to use but dare not use" accurately describes the mentality of many U.S. enterprises. Microsoft is also evaluating the possibility of letting DeepSeek take over part of the workload of Copilot, which was specifically mentioned as evidence in Cotton's letter.

So far, only the U.S. has blocked the use of Chinese models. If we look at Beijing, another hand is also reaching for the kill switch.

The Ministry of Commerce of China is negotiating with Alibaba, ByteDance, Zhipu AI and several other companies to include model weights into export control. This policy has not yet been implemented, but the general direction is very clear.

At the same time, the new EU regulations have come into effect. About 190 institutions have signed the transparency code of conduct, including Anthropic, OpenAI, Google and Microsoft, but no Chinese large model company has signed it.

Some outsiders interpret this as arrogance, but the more reasonable explanation is that the relevant control measures are also under preparation in China, and no enterprise dares to promise to keep the door open permanently.

The two kill switches on both sides are turning in the same direction unexpectedly.

Politicians are still writing letters, but the market has already made deals. On September 3, NVIDIA officially announced the acquisition of Hugging Face for 12.93 billion U.S. dollars. 11.9 billion of the total consideration is cash, and the remaining 1 billion is reserved for core employees. The transaction is subject to antitrust review and is expected to be completed in the first half of 2027.

Hugging Face is the hub of AI open-source models and the first stop for Chinese models to go global. ChatGLM, Qwen and DeepSeek all went to the whole world from this platform.

The platform changed its owner overnight. NVIDIA claimed that it would remain neutral and allow everyone to upload and download models, but its actions tell a different story. As early as March this year, it took the lead in forming an open-source alliance called Nemotron, bringing in Cursor and Thinking Machines to develop the U.S.'s own open-source models.

The irony is that Cursor uses Kimi, and the architecture of Thinking Machines follows the path of DeepSeek. The U.S.'s counterattack in the open-source field has already begun, but it is still based on the technical foundation from China.

After providing free services for such a long time, someone has to start charging. The reasons are also very convincing: the number of users has grown too fast, the computing power is insufficient, the cost of training and electricity bills is real, and shareholders are waiting for returns. The free model cannot be sustained for much longer.

Kimi is currently negotiating with Microsoft, Google and Amazon to sell its models on their cloud platforms, charging up to 30% of the revenue share. Alibaba is also considering charging the heavy users of its models.

Zhipu AI has taken the third path. It has set a threshold in its new agreement: large cloud vendors with annual revenue exceeding 10 billion U.S. dollars must pass a security review before using its models. It does not charge fees, but requires identity registration.

DeepSeek is still providing its models for free under the most permissive open-source agreement, without taking any share of the revenue. Among the vendors that used to distribute free leaflets together, some have started to charge franchise fees. The weapon of free is being half-dismantled by its own practitioners.

Ethan Mollick, a professor at the Wharton School, has a farsighted view. He says the era of cutting-edge open-weight models may not last much longer.

Both sides can understand this sentence in their own way: Washington thinks the era lasts too long, while Beijing thinks it is too short.

For the models that have already been downloaded to local