HomeArticle

Overseas open-source models are accelerating their development pace again: Mistral and the "US version of DeepSeek" show their hands simultaneously.

极客邦科技InfoQ2026-10-08 11:03
The Western open model camp has begun to accelerate.

The Western open-model camp has started to speed up its pace.

On local time October 5, Reflection, a U.S. startup known as the "U.S. version of DeepSeek", released Beam, its first open-weight model; one day later, France's Mistral launched Mistral Large 4, its largest model to date. The two companies chose highly consistent main comparison targets when releasing their models: GLM, Kimi, DeepSeek, Qwen.

Data from Hugging Face shows that in the past year, Chinese models have accounted for about 41% of the total downloads on the platform, with both monthly downloads and cumulative downloads exceeding those of the United States. Models including DeepSeek, Alibaba Qwen, Moonshot AI Kimi, Zhipu AI GLM, MiniMax, and Xiaomi MiMo have continuously refreshed the capability ceiling of open models in turn, making Chinese open models the target that global open-source developers are catching up with.

On the other hand, open models are no longer "niche alternatives" only used by researchers and geeks: production traffic data from Vercel AI Gateway shows that open-weight models processed 56% of the tokens on its platform in August this year, while this proportion was less than 10% last December. Therefore, the global open model competition is inevitable.

U.S. and European Open Models Debut Simultaneously

The "U.S. version of DeepSeek" submits its first answer sheet: there is still a gap

Highly anticipated Reflection was founded in 2024, and both of its founders are from Google DeepMind: CEO Misha Laskin was once in charge of reward modeling for Gemini, and CTO Ioannis Antonoglou participated in the development of AlphaGo and AlphaZero, and later led the reinforcement learning from human feedback (RLHF) work for Gemini.

However, Reflection did not initially set out to become the "U.S. version of DeepSeek". The company initially focused on autonomous programming, hoping to use reinforcement learning to build a "superintelligent system" that can independently complete complex tasks. At the same time, it originally judged that sufficiently powerful open models would continue to emerge in the West, and Reflection could build directly on top of these models.

It was DeepSeek V3 that truly changed the company's route. According to the founders of Reflection, after Chinese open models advanced rapidly while no equivalent alternatives appeared in the West for a long time, the company decided to train cutting-edge open models on its own, and gradually clarified its goal as "bringing the open model frontier back to the West".

Two years later, Beam is considered Reflection's first real submission of its answer sheet.

Beam is a Mixture-of-Experts (MoE) model with a total parameter count of 501 billion, activating 23 billion parameters per token, and is mainly oriented to programming, reasoning and Agent workloads. The model was pre-trained with 23.8 trillion tokens, and Reflection deployed 10,500 Nvidia GB300 GPUs to run 4 consecutive weeks of high-computing-power reinforcement learning, generating more than 100 million rollouts. The model is currently in the final red-teaming test phase, and Reflection plans to release the model weights, technical report and development tools in October.

According to the test results released by Reflection itself, the biggest highlight of Beam is not its absolute capability, but its reasoning efficiency. The company claims that on some advanced reasoning tasks, it can achieve comparable performance with about 1/4 to 1/3 of the reasoning computing power of GLM-5.2.

However, when compared horizontally with China's latest generation of models, there is still an obvious gap for Beam to truly "catch up".

For example, in the Agent programming test DeepSWE v1.1, Beam scored 44.4, which is almost on a par with GLM-5.2's 44.0, but lower than GLM-5.3's 61.0, Kimi K3's 68.0, and DeepSeek V4.1 Flash's 74.2; on Terminal Bench v2.1, Beam scored 80.1, while GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash reached 88.2, 88.3 and 90.6 respectively. On Humanity's Last Exam (HLE), Beam scored 36.2, which is also lower than GLM-5.3's 42.3 and Kimi K3's 46.9.

Therefore, currently Beam has pulled U.S. open models back to the vicinity of China's top models, but it has only caught up with the previous-generation GLM-5.2, and there is still a gap of about one generation with China's latest round of open models.

Reflection itself does not shy away from this. The official directly admits that Kimi K3 is still leading in absolute capability, and shifts Beam's selling point to "how much intelligence can be obtained per unit of reasoning computing power". CEO Misha Laskin positions Beam as a "mainstay model" that enterprises can run for a long time, rather than a maximum-sized model that only pursues rankings.

This is very similar to part of the logic that DeepSeek initially impressed the market: the value does not lie in whether the model ranks first, but how much it costs to reach a certain level of capability. The difference is that Reflection can hardly be called a "low-cost version of DeepSeek".

For Beam, the last round of reinforcement learning alone occupied more than 10,000 of the latest GB300 GPUs at the same time; the company has raised billions of dollars in financing, obtained large-scale investment from Nvidia, and locked in a large amount of computing power through partnerships with SpaceX, Nebius and other parties. It is a "heavy-asset version of DeepSeek" backed by U.S. capital and Nvidia's latest hardware.

Ioannis said that the company's goal in the next few years is to eventually make the "open frontier" and the "closed frontier" the same frontier. In 2027, Reflection plans to release larger-scale models, continue to expand reinforcement learning computing power, and launch multiple reinforcement learning enhanced versions in succession on top of the base model, just like models such as GLM.

Misha believes that behind this route is actually a larger judgment: reinforcement learning has evolved from a domain-specific technology in AlphaGo and AlphaZero in the past to a basic method that can continuously promote the progress of general-purpose models, programming models and Agent models.

In his view, these models are "either already some form of AGI, or at least clearly on a continuous path leading to AGI". The core issue next is no longer just how to continue to strengthen the model, but how to engineer and productize this capability, and spread it through a multi-polar, more open ecosystem.

Mistral re-sprints towards the top tier of open models

Compared with Reflection, which has just launched its first base model, European model vendor Mistral is facing even greater pressure.

Two or three years ago, it was once the most representative company in the open model field, but as models such as DeepSeek, Qwen, Kimi and GLM iterated rapidly, Mistral's presence on the frontier capability rankings continued to decline. In September this year, CEO Arthur Mensch was even directly asked in an interview: Mistral's models have fallen out of the top 20 in the Artificial Analysis ranking, is it left behind by the United States and China?

Mensch responded at that time that the new generation of large models is about to be released, and the company has planned the next two generations of products for the coming year; a large part of the newly raised financing of Mistral will be used to expand training computing power, and the amount of training computing that can be called in the future may increase by about 20 times.

Two weeks later, Mistral Large 4 was unveiled.

Mistral Large 4 has 52 billion activated parameters and 1.05 trillion total parameters, equipped with a 1.6-billion-parameter visual encoder, natively supports multimodality and more than 160 languages, and its context window reaches 1 million tokens. It is worth noting that it was trained from scratch in Mistral's data centers in Europe using about 3,800 Nvidia Grace Blackwell GPUs; the model has now opened API preview, and the full weights will be released at the end of October.

According to the Coding Agent Index officially announced by Mistral, Large 4 has surpassed DeepSeek V4 Pro and Qwen 3.8 Max, but is lower than Kimi K3; in the blind test conducted in cooperation with third-party Surge AI, professional reviewers gave Large 4 a code quality score of 3.74, which is lower than Claude Opus 5's 4.22, but higher than Kimi K3's 3.59, GLM-5.3's 3.60 and GLM-5.2's 3.40.

Mistral is obviously targeting "enterprise Agents" rather than just developers. In 657 cross-application business workflows across Gmail, Google Sheets, Slack, Salesforce and other applications in AutomationBench, Mistral claims that Large 4 surpasses Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro, but is lower than GLM-5.3. On scientific coding, cybersecurity, and some financial and legal tasks, Mistral also claims that it has reached the top level of open models.

This time, Mistral also specially introduced its cybersecurity capabilities. Mistral says that Large 4 has entered the top 5 globally in the Artificial Analysis Cyber Index, reaching 82% on one task of "reproducing and fixing real vulnerabilities", which is the highest score in this test; it completes 93% of the challenges on Cybench. Mistral also deliberately emphasized that some closed-source models can hardly answer some vulnerability reproduction tasks due to their safety refusal mechanism, while open-weight models allow enterprises to define their own security policies and deploy them in private clouds or on-premises environments.

Of course, a large part of these data currently comes from Mistral's own tests, and Large 4 is still in the public preview phase, with the real weights not yet released.

Complementing the business model that Chinese vendors have initially formed a prototype of

Apart from model capabilities, both Mistral and Reflection are also following a business route that has been initially verified by domestic vendors.

Reflection tells a very "American-style" business story: it hopes to build a new full-stack AI infrastructure provider, starting from base models, undertaking reasoning, Agents, software and enterprise deployment upwards, delving deep into computing power management and cluster scheduling downwards, and even owning its own data centers in the long run.

He believes that every layer including models, applications, cloud computing, and bare metal will be commoditized in the long run, and the companies that can continuously obtain high profits are often those that vertically integrate the entire technology stack. Misha even gave a very "Silicon Valley-style" formula:

Revenue potential = intelligence density × computing power scale × enterprise trust.

Misha said that open models themselves do not necessarily make money directly, but they can create huge demand for reasoning. Just like Kubernetes is open source itself but promotes the cloud market, open models will drive enterprises to purchase computing power, reasoning and complete AI infrastructure. Therefore, what Reflection really sells is the capability to build a complete set of intelligent infrastructure around Beam.

It even compares itself to early AWS and GCP. AWS was initially developed by Amazon to solve its own server infrastructure problems, and Google Cloud came from Google's large-scale distributed system capability formed to support its search business. The logic of Reflection is that in order to train cutting-edge models, it originally had to solve problems such as serving trillions of tokens, billions of Agent sandboxes, and GPU scheduling, and these internal technologies can eventually be exported outward just like cloud computing.

Mistral's business model also emphasizes enterprise deployment, but it is more biased towards a "European version of full-stack AI".

Arthur advocates that small, efficient open models that can be customized and deployed are more in line with enterprise needs than general-purpose models that only pursue maximum scale. Mistral is building its own models, enterprise customization platforms and European computing power infrastructure at the same time, hoping to become a supplier whose entire process from training, reasoning to deployment can happen within the European legal and infrastructure system.

In his view, European enterprises cannot completely build their core AI capabilities on whether foreign vendors are willing to continue to provide services, nor can they hand over their own technology base to models from other countries. Therefore, Mistral not only trains its own models, but also builds data centers in Europe, hoping to control the entire chain from computing power, models to enterprise deployment.

Attitudes towards domestic open models

One of the two companies is located in Europe, and the other comes from the United States, but the recent public statements of their founders show that they are forming an increasingly similar judgment on open models: the AI market has moved from the stage of "who can provide the most usable API" to the stage of "who owns the models and who controls the infrastructure".

On the question of "why must we build open models", Reflection's explanation is: the AI market is moving from "renting intelligence" to "owning intelligence".

Misha compares closed-source APIs to "renting a house". In the past three years, the AI commercial market has just taken shape, and enterprises calling models from companies such as OpenAI, Anthropic and Google are essentially renting intelligence. This model is simple enough and efficient enough in the early stage of the market. But as AI truly enters the core business of enterprises, things begin to change.

"Startups that were still very small a few years ago have grown into large enterprises, and large enterprises themselves have begun to truly use AI, so the ownership market naturally emerges." Misha said that if a company wants to truly own intelligence, it must be able to deploy the model on its own infrastructure, control data, modify the model, and carry out customization, which means the model must be sufficiently open.

Ioannis said from a technical perspective that the gap between the most cutting-edge open models and closed models is narrowing, and "there is no technical reason that determines that open models must always lag behind". In his view, as long as there are sufficient talents and computing power, open models are fully capable of reaching the same level as closed models. In the past, enterprises did not use open models, and one core reason was only that "the capability was not enough"; now, a number of Chinese open models can directly replace closed models to complete a large number of tasks, and the premise for commercial migration has been established.

Mistral CEO Arthur Mensch's judgment is very close to Reflection's on this point. In his view, as the capability of open models improves, a large number of enterprises no longer need to continue to pay additional API premiums for closed models. He recently said that about 99% of the enterprise scenarios that Mistral has contacted can be completed with open models.

So why did Chinese models come out on top first in the past year?

The two companies also have quite consistent judgments on this: the leading position of Chinese open models is not just a coincidence that a single company launched a good model, but the result of the combined effect of commercial incentives, industrial environment and training resources.

Misha said that the United States actually used to have the most powerful open models, and Llama 3 is an obvious node. But at that time, the U.S. AI commercial market was still dominated by APIs, and Meta or other companies did not have enough strong commercial motivation to continuously open the most cutting-edge capabilities.

The situation in China is different. DeepSeek V3 is an important turning point. For the first time, open "intelligence" from China has formed a large-scale impact globally, after which a more concentrated advancement of open models emerged in China. Misha believes that a large part of the motivation behind this comes from the fact that building one's own open-weight ecosystem has industrial and strategic value in itself.

"In fact, these Chinese labs have not made much money until recently. It is only now that their model capabilities are strong enough, and the market has matured enough to pay for open intelligence, so they have begun to generate real revenue," Misha said.

Arthur explained from a more engineering perspective: although the computing power of China's top labs is lower than that of the several largest AI companies in the United States, it is still higher than that of Mistral; at the same time, they will make the most of the data on the Internet, and their iteration methods have been industrialized. These conditions allow Chinese labs to iterate rapidly.

Will open models be more dangerous?

Although both companies are actively defending open models, neither of them describes themselves as extreme openists who believe that