HomeArticle

China's Model, Global Foundation

字母榜2026-09-07 18:34
China's large AI models are quietly becoming the "water, electricity and coal" — the indispensable foundational utilities — of the global AI industry.

It is no longer news that Chinese large models are being "shelled" by overseas players, and a new case just broke out a few days ago.

On September 2, Saudi AI company HUMAIN launched the HUMAIN M3 model. At the initial release, the official directly claimed that it was a "100% Saudi Model".

The name of this company may sound unfamiliar to most people, but the term "M3" looks extremely familiar, as MiniMax released its MiniMax M3 not long ago.

In addition, its configuration of 428B total parameters and 23B activated parameters is exactly the same as that of MiniMax M3. Soon, many people in the comment section raised doubts: Is this really 100% Saudi?

Sultan Alfaifi, the founder of HUMAIN, later had to clarify in person that HUMAIN M3 was developed based on MiniMax's M3.

However, this incident is not an isolated case in the Chinese AI industry. Japan's Rakuten uses DeepSeek, Singapore's national AI initiative uses Qwen... Chinese models have quietly become the first choice for AI developers around the world.

In the past, when Chinese large models went global, they mostly sold APIs and products, earning revenue from invocation fees and service fees. Now, they are expanding into the global market in a brand new form.

Leaving aside the pros and cons for now, there is no doubt that this is a signal that open-source models have truly begun to spread outward on a large scale.

1

HUMAIN actually has its own self-developed model.

In August 2025, this national-level AI company established by the Saudi Public Investment Fund just released the ALLaM 34B model, with more than 120 AI experts and 35 doctors participating in the development. The model was trained from scratch in Saudi Arabia, focusing on alignment with the Arabic language and local culture, and the official emphasized at that time that it was "built from the ground up".

However, when it came to launching its flagship M3 model in September 2026, Saudi Arabia still chose a Chinese model as the base, and only used its own corpus for post-training.

As mentioned earlier, it is not new for countries around the world to use Chinese open-source models.

On March 17, Japan's Rakuten Group grandly released Rakuten AI 3.0, which was officially described as "the largest high-performance AI model in Japan". The launch was backed by GENIAC, a national Japanese AI project jointly promoted by the Ministry of Economy, Trade and Industry and NEDO, and Hiroshi Mikitani, founder, chairman and CEO of Rakuten Group, known as "Japan's Jack Ma", personally endorsed the product.

However, just a few hours after the release, developers found that the "model_type" field in the configuration file of Rakuten AI 3.0 was marked as "deepseek_v3".

What is more embarrassing is that Rakuten directly deleted DeepSeek's MIT license file in its initial upload, and replaced it with its own Apache 2.0 license. It is equivalent to taking something that others have given to everyone for free, pasting its own label on it and selling it at a clear price.

After being exposed by the community, Rakuten hurriedly added a NOTICE file to admit that the copyright belongs to DeepSeek.

In contrast, Southeast Asian countries are much more transparent.

Singapore's national AI initiative AISG previously used Meta's Llama as the base model for training in its SEA-LION project, and later switched to Alibaba Cloud's Qwen3-32B as the base to launch Qwen-SEA-LION-v4.

According to Alibaba Cloud, AISG injected more than 100 billion Southeast Asian language tokens into this base for post-training, and the model ranks first among open-source models on the SEA-HELM Southeast Asian language leaderboard.

In May 2025, the Malaysian Ministry of Communications launched a project known as Southeast Asia's first "sovereign AI infrastructure", which uses DeepSeek as its base model. This is also DeepSeek's first overseas deployment at a national scale.

India is the most interesting control group, as the two cases there are completely different in nature.

The 32B lightweight model Alpie launched by 169PI was called "India's version of DeepSeek" by Indian media. However, developers later found through the development documentation that it was no wonder it was called India's DeepSeek, as it was essentially DeepSeek R1.

The full name of this model is DeepSeek-R1-Distill-Qwen-32B, which means that DeepSeek R1 is used as the teacher model to output high-quality reasoning results, which are then distilled to a 32B student model with Qwen as the base.

There is another company called Sarvam AI, which received a large amount of computing power subsidies from the IndiaAI Mission, about 990 million rupees, equivalent to about 11 million US dollars. The official claimed that the model was "trained from scratch", but after the community checked its configuration, they found that the multi-head latent attention, GRPO reinforcement learning and other mechanisms all came from DeepSeek.

Subsequently, the CEO of Sarvam AI also admitted that the team "admires and learns from DeepSeek".

The next country is South Korea. According to foreign media reports, at least 3 out of 5 finalist teams in its national large model competition, the Sovereign AI Foundation Model Project, used Chinese open-source models.

For example, the visual and audio encoders of Naver's team share similar features with Alibaba's Qwen, the reasoning code of SK Telecom's team is identical to DeepSeek, and some components of Upstage's team were directly copied and pasted from Zhipu's GLM. The investigation also found that Upstage's team even failed to delete Zhipu's copyright mark completely from its project.

Similar incidents have also taken place in Europe. Cursor, which was acquired by Elon Musk not long ago, launched its programming model Composer 2 in March. However, developers found from the API response that the actual ID of the model was "kimi-k2p5-rl-0317-s515-fast".

This forced Cursor to admit that Composer 2 was built based on Moonshot AI's Kimi K2.5, through an authorized cooperation with Fireworks AI.

Moonshot AI is very generous, and the official later posted a congratulatory message, stating that "Kimi K2.5 provides base support for Composer 2".

Thinking Machines, founded by OpenAI's former CTO Mira Murati, is another example of this kind.

Its 975B open-source model Inkling launched in July clearly states in the official technical documentation that its MoE architecture is derived from DeepSeek-V3.

In addition, Siemens, Renault and Orange all publicly announced at VivaTech that they have adopted a "China-US-EU hybrid model" strategy.

Chinese models have quietly become the base of global AI development.

2

The most obvious advantage is low cost.

If you train your own large model from scratch, the cost of the data center alone can easily reach tens of millions of dollars.

But models like DeepSeek and Qwen have open weights and are licensed under MIT, which allows commercial use, weight modification and local deployment. Although Meta's Llama is also open source, it does not allow large-scale use, so national institutions and large enterprises cannot adopt Llama.

Moreover, secondary development open-source tools such as ms-swift are complete, with native deep adaptation to Qwen, DeepSeek and GLM. All corresponding functions including LoRA, QLoRA, DPO, GRPO, quantization and distributed training are well encapsulated, so that fine-tuning and alignment can be completed with just one command.

To put it simply, it is like pre-made dishes with seasoning packets, which you can just heat up in the microwave and serve.

The MoE architecture adopted by DeepSeek allows each token to activate only 37B parameters, which greatly reduces the actual reasoning cost, and the operating cost of the entire model will drop accordingly after deployment.

Chinese open-source models have far fewer restrictive rules.

Countries have very tangled attitudes towards AI development: they want to develop AI, but they are afraid of data going out of their borders, and they are also worried that the access to closed-source APIs will be cut off.

In June 2026, the White House issued an export control order to Anthropic, requiring it to suspend foreign users' access to its most advanced Fable 5 and Mythos 5 models. After receiving the order, Anthropic announced that it would disable these two models.

Two weeks after the ban was issued, according to data from 4SAPI, 68.4% of medium and large enterprise customers reported reduced efficiency of their core AI workflows, and 31.6% of small and micro developers migrated part of their workload to other models, with the average reasoning cost rising by 27.3% during the migration period.

That's not all. After issuing the disablement ban, Anthropic only gave all users 90 minutes to complete the migration. It is easy to imagine how desperate users outside the United States were at that time.

European enterprises are particularly sensitive to this. The license of Llama 4 even explicitly denies EU developers the right to use its multimodal capabilities, while the MIT and Apache 2.0 licenses adopted by Chinese models do not have such restrictive clauses.

As a result, "local deployable, weight accessible" has become a rigid demand, and the Chinese models represented by DeepSeek and Qwen in the open-source camp are exactly the ones that can meet these requirements at the same time.

More importantly, Chinese models perfectly match the actual needs of other countries.

Countries that use open-source models often have demand for small languages. The pre-training of the Qwen series covers 119 languages and dialects, which is inherently suitable for non-English markets such as Southeast Asia and the Middle East.

At the same time, DeepSeek and MiniMax M3 have strong reasoning and tool calling capabilities, which meet the secondary development needs of various countries to build "local national models + local Agent applications".

Conversely, this is also a subtle form of Chinese-style outward influence.

Models from closed-source vendors are commercial-oriented and emphasize paid services. But Chinese vendors' attitude is more like that of GitHub, who are more willing to cooperate with others.

For example, Moonshot AI can authorize Kimi to Cursor through a third-party partnership. This "non-competing" gesture is more effective than any marketing strategy in the current era where sovereign narrative is dominant.

Therefore, rather than saying that Chinese models have become the base of the whole world, it is more accurate to say that they have been "chosen" by the global market.

3

In the United States, being used as a model base is a clearly priced business.

On January 12, 2026, Apple and Google officially announced a multi-year cooperation, under which Gemini will power the next generation of Siri and Apple's base model.

The price has not been made public, and foreign media estimate that the contract is worth about 1 billion US dollars per year.

In fact, Apple and Google are strong competitors, but in order to equip its voice assistant with a powerful "brain", Apple still chooses to pay 1 billion US dollars a year to its competitor.

There is also cooperation between Samsung and Google. Although there is no direct cash transaction, the two sides have a deep binding relationship.

The core model of Samsung's Galaxy AI also chose Gemini, covering about 400 million devices in 2025, with a target of doubling to 800 million devices in 2026. In return, Samsung gives Google sales revenue share, and pre-installs the full suite of Google apps as the default entry on its mobile phones.

The cooperation between Amazon and Anthropic follows a different model.

Amazon has cumulatively invested 8 billion US dollars in Anthropic, and announced in April 2026 that it will invest another 5 billion US dollars first, and up to 20 billion US dollars more depending on the results of Anthropic, with a potential total investment of up to 33 billion US dollars.

In exchange, Anthropic promises to spend more than 1 trillion US dollars on AWS in the next 10 years, and run the training of its main models on Amazon's self-developed Trainium chips.

Microsoft and OpenAI also have such a deep binding relationship. Microsoft has cumulatively invested more than 13 billion US dollars. According to the agreement, Microsoft will first take 75% of the profit share until it recovers its investment, and then the share will drop to 49%. On April 27, 2026, the two sides further announced the termination of their exclusive cooperation, and set an upper limit on the payment of profit sharing.

Products including Bing, Edge, M365 Copilot and GitHub Copilot are all built on top of GPT.

Chinese open-source vendors do not seem to be very interested in this base model business, and they pay more attention to the compound interest effect brought by the ecosystem and standards.

Chinese vendors give away the model weights for free: no license fees, no profit sharing, and they even cooperate with their partners to package the model as a source of national pride for local people. So the question arises: if they don't charge money, what are they pursuing?

The more countries adopt a model, the better its multilingual capabilities will be improved. Different national conditions will lead to different usage scenarios of the model, which will also improve the stability of its Agent applications.

When more and more countries build their national models on Chinese base models, Chinese models will become the de facto standard. For example, in the field of multimodal models, Qwen-VL has long become the industry standard.

According to Presenc AI's statistics on the total download volume of Hugging Face, Qwen3-VL-2B-Instruct is the second most downloaded model on the entire site, second only to the 80MB embedding model all-MiniLM-L6-v2, and it is also one of the only two models on Hugging Face that have exceeded 100 million downloads.

The value of standards far exceeds that of license fees.

This article is from the WeChat official account "Letter List" (ID: wujicaijing), written by Miao Zheng, and republished with authorization from 36Kr.