HomeArticle

Everyone claims to be the number one, who on earth is the most powerful large model in China?

砺石商业评论2026-09-01 11:47
The first echelon of large language models in China is now a duopoly dominated by Kimi K3 and Qwen3.8-Max.

A relatively honest statement is that the first tier of China's large model industry presents a two-strong pattern consisting of Kimi K3 and Qwen3.8-Max. The former excels in international verification, while the latter takes the lead in comprehensive Chinese capability and ecosystem.

At present, a peculiar phenomenon has emerged in China's large model industry: almost every leading company claims that its model is "No.1 in China" or even "Top 3 in the world".

Alibaba states that Qwen3.8-Max "ranks second only to Claude in third-party blind tests"; Moonshot AI claims that Kimi K3 is "the world's strongest open-source model"; Zhipu says GLM-5.2 "ranks No.1 globally on Code Arena"; DeepSeek's model card marks "SWE-bench Verified 80.6%"; ByteDance's Doubao 2.1 asserts that it "leads the world in IMO gold medal standards and HLE benchmarks"; Baidu's Ernie 5.0 emphasizes "10M context window, No.1 in Chinese writing".

All these claims are "true", but none of them are complete. They share one common feature: every "No.1" is preceded by a carefully chosen attributive. The model with the largest parameter size can be called No.1, the top scorer in a single benchmark can be called No.1, the leading entry on a certain blind test list can be called No.1, and the best performer in a certain price range can be called No.1. With different attributives, the conclusions naturally do not conflict with each other. It is just like the smartphone era when every manufacturer claimed their product "ranks No.1 in running scores, No.1 in camera performance, No.1 in battery life", and similar to the new energy vehicle industry where every brand asserts their car "leads in range, intelligent driving and safety".

The situation on the other side of the ocean is different. There is a relatively recognized landscape in the US large model industry: OpenAI takes the lead in certain stages, Anthropic's Claude leads in other stages, and Google's Gemini overtakes alternately in between. The consensus by mid-2026 is that Anthropic is back at the top, with Claude Fable 5 ranking first on the Artificial Analysis Intelligence Index at a score of around 60, and GPT-5.6 Sol ranking second at around 59. This consensus is formed through joint voting by third-party evaluation institutions, developer communities and real usage data, without the need to refer to any company's press conference.

China has not formed such a consensus. It is not that there is no answer, but that the answer is covered by the noise of public relations promotion. This prompts us to set aside all self-claimed statements from companies, and mainly adopt three most authoritative types of third-party verification data — Arena global blind test, Artificial Analysis Intelligence Index and SuperCLUE Chinese evaluation — for cross-reference, to see what conclusion the facts actually support.

First chain of evidence: Artificial Analysis Intelligence Index

Artificial Analysis (AA) is the most cited independent evaluation institution in the international developer community. Its Intelligence Index integrates standardized tests of multiple dimensions including mathematics, reasoning, coding and knowledge, ranking 189 models worldwide on the same track. It does not accept funding from manufacturers, and adopts a unified testing standard, making it the closest thing to an "objective" third-party measurement at present.

As of early August 2026, the scores of major Chinese models on the AA Intelligence Index are as follows:

Moonshot AI Kimi K3 (Max): 57 points, ranking 4th among 189 models. This is the highest ranking that any open-source weight model has ever achieved, surpassing Claude Opus 4.8 (around 56 points), second only to Claude Fable 5 (around 60 points), GPT-5.6 Sol (around 59 points) and a flagship model from Google.

Zhipu GLM-5.2 (Max): 51 points, ranking outside the top 10.

DeepSeek V4 Flash 0731: 50 points, on a par with Google Gemini 3.6 Flash.

Alibaba Qwen3.8-Max: No score available yet. As of now, AA has not completed its evaluation. The verified score of its previous generation Qwen3.7-Max is 56.6 points — this is the only third-party reference for this model family.

This table reveals two facts. First, China's strongest model has reached the threshold of the global first tier: the 3-point gap between Kimi K3's 57 points and the top score has narrowed from a double-digit gap a year ago. Second, the "officially claimed strongest" and the "verified strongest" are two different things — Qwen3.8-Max, which claims to be "second only to Claude Fable 5", is precisely the only leading model that has no third-party score yet.

Second chain of evidence: Arena blind test

The mechanism of Arena (arena.ai) is human blind test: two anonymous models answer the same question, and users vote for the better one. Millions of votes are converted into ELO points. It does not measure "how many points a model scores in exams", but "which one real people find more user-friendly". On the list of Week 33 of 2026 (August 10-16), the results of domestic models are as follows.

In the overall ranking, Anthropic's Claude series takes the top 5 spots. Among domestic models, Qwen3.8-Max ranks 8th (ELO 1491), and Kimi K3 Max ranks 12th (1489), which are the only two Chinese models that enter the top 15. This pattern is worth a closer look: the score difference between the two is only 2 points, which means they are basically tied within the error range, but Qwen3.8-Max ranks higher.

In the coding ranking, Kimi K3 Max ranks 6th (1544 points), which is the highest-ranked domestic coding model; Qwen3.8-Max ranks 13th, and Xiaomi Mimo ranks 25th.

The front-end development ranking is the list where domestic models deliver the most outstanding performance. Kimi K3 Max ranks 2nd with 1674 points, only 18 points behind the leading Claude Opus 5 Max (1692 points); Qwen3.8-Max ranks 3rd (1667 points); Zhipu GLM-5.2-Max ranks 9th; DeepSeek V4 Pro ranks 10th right after its new entry. When Kimi K3 was first released a month ago, it briefly topped this list with 1679 points, which was the first time a domestic model got the global No.1 spot on Arena's practical ranking.

In the agent ranking, Kimi K3 Max ranks 5th with a net improvement rate of 10.60%, standing in the global first tier together with the Claude Opus series; Qwen3.8 Max ranks 11th after its new entry; GLM 5.2 Max ranks 14th; DeepSeek V4 Flash ranks 19th.

Putting the four lists together, the conclusion is clear. In the Arena system, Kimi K3 Max is the Chinese model with the widest coverage and highest position — 6th in coding, 2nd in front-end development, 5th in agents, 12th in overall ranking, none of its rankings fall out of the top 15. Qwen3.8-Max is the domestic model with the highest position on the overall ranking (8th), but lags behind Kimi on the two practical rankings of coding and agents.

Third chain of evidence: SuperCLUE Chinese evaluation

SuperCLUE is one of the most credible third-party benchmarks in Chinese scenarios. On the July 2026 list, the overall total score ranking of 12 mainstream domestic models is: Qwen3.8-Max, Kimi-K3, Doubao-Seed-2.1-Pro and DeepSeek-V4-Flash.

The score gap between the first and the second is very small, but the order is clear. In terms of comprehensive Chinese capability, Alibaba's Qwen3.8-Max ranks first. This echoes its performance on Arena's overall ranking (8th, higher than Kimi's 12th place).

It is worth noting the third and fourth places. ByteDance's Doubao 2.1-Pro ranks third in the comprehensive Chinese evaluation, and it is the only model in the "top 5" that has not participated in the Arena international blind test. Its capability is visible on domestic lists, but missing in the international coordinate system. DeepSeek ranks fourth, and its former leading position has been taken over by latecomers.

In the specialized programming evaluation, SuperCLUE gives another result. GLM-5.2 leads by a large margin with 64.13 points, and Kimi K2.7-Code scores 55.43 points. Zhipu is a typical "specialized student": it ranks No.1 globally in a single programming category (it has also taken the 1st spot on the international Code Arena and Design Arena), but its comprehensive capability cannot rank in the top 4 in China.

Portraits of the five top players: their respective strengths

Before drawing the portrait of each model, we first put the hard specifications of the five flagship models together, because the summer of 2026 is the first summer when the five companies intensively released cutting-edge models within six weeks: Zhipu GLM-5.2 in mid-June, Kimi K3 on July 16, Qwen3.8-Max preview on July 19, open-source Kimi K3 on July 27, DeepSeek V4 Flash 0731 on July 31, and the official version of Qwen3.8-Max on August 3.

Three facts are worth noting. First, the million-token context window has become a standard configuration for flagship models, no longer a differentiator — all five models support it. Second, all 2-trillion-parameter-level models choose to go open-source (under MIT license), which is a collective contribution of Chinese companies to the global open-source community, as well as a differentiated competition against the US closed-source route. Third, as flagship models, their output prices differ by 50 times — this price gap itself is a choice of business model: Kimi K3 matches the highest reasoning density with the highest price, while DeepSeek captures the high-frequency Agent calling market with an extremely low floor price.

By cross-referencing the three chains of evidence, the actual positions of China's top 5 large models in August 2026 can be described as follows.

Moonshot AI Kimi K3, China's No.1 in the verification dimension. Released on July 16, open-sourced on July 27, with 2.8 trillion parameters and 104B activation, it is the open-source model with the largest parameter scale in human history, as well as the world's first trillion-level open-source model that natively supports four modalities: text, image, audio and video. Its third-party report card is the most complete among all domestic models: 57 points on the AA Intelligence Index (the highest in open-source history, 4th globally, surpassing Claude Opus 4.8), 2nd place on Arena front-end ranking, 6th on coding ranking, 5th on agent ranking, and 2nd on SuperCLUE Chinese evaluation. Actual tests by the third-party evaluation institution FireworksAI on 1030 real Agent tasks show that Kimi K3's performance is close to Claude Fable 5, while its cost is about 1/50 of the latter. Its shortcoming is the high price: the output price is 100 RMB per million tokens, 50 times that of DeepSeek V4 Flash.

Alibaba Qwen3.8-Max, No.1 in comprehensive Chinese capability and ecosystem. With 2.4 trillion parameters and 95B activation, its official version was released on August 3. It ranks first in SuperCLUE comprehensive Chinese evaluation, and is the highest-ranked domestic model on Arena overall ranking (8th). But it must be pointed out that its international verification report card is currently blank. It has not been evaluated by the AA Intelligence Index, and all its official benchmarks (4.16 times return in e-commerce simulation, 265 independent programming submissions in 16 days, etc.) come from Alibaba itself. Its real moat lies elsewhere: the Qwen open-source family has accumulated the world's largest download volume and the most abundant derivative model ecosystem over the years, which are assets that the other four companies do not have.

DeepSeek V4, the king of reasoning and cost performance, but has been overtaken. In early 2025, DeepSeek R1 approached OpenAI's top model at an extremely low training cost, causing a stir in Silicon Valley and a sharp drop in Nvidia's stock price on the same day, making the world pay attention to China's large models for the first time. That was the highlight moment for DeepSeek, and also the highlight moment for China's large model industry. In the following year and a half, DeepSeek slowed down its pace: V4 was released on April 24 this year, with no heavy updates in the four months in between, and the updates of V4 Flash 0731 on July 31 and V4 Pro in mid-August brought it back to the competition. Its foundation is still solid — in an independent horizontal evaluation (actual API test) of 8 mainstream models in mid-August, DeepSeek V4 Pro ranked first among all domestic models with 9.1 points, second only to Claude Opus 4.8 (9.20) across all models. It got 9.5 points in the reasoning category, surpassing GPT-5.6, and its capability of "identifying insufficient conditions and knowing when not to draw a conclusion" was rated as the strongest across all models. Its SWE-bench Verified score of around 80.6% is the highest verifiable standardized score among domestic models. V4 Flash (284B/13B activation) reduces the output price to 2 RMB per million tokens, about 1% of that of Claude Opus, making it the default option for high-frequency Agent calling scenarios. But its AA Intelligence Index is 50 points, and its overall ranking has been left behind by Kimi K3 (57 points) and Qwen3.8-Max. It has changed from the global leader in 2025 to a chaser in 2026. This story itself is a footnote to the fierce competition in China's large model industry: in this industry, stopping iteration for a quarter may make you fall out of the first tier.

Zhipu GLM-5.2, a programming specialist. With about 744B parameters, open-sourced under MIT license, and trained on Huawei Ascend chips, this has special strategic significance in the context of 2026. It has taken the global No.1 spot on two international programming blind test lists, Code Arena and Design Arena, and leads by a large margin in the SuperCLUE specialized programming evaluation. But its comprehensive capability (51 points on AA, 14th on Arena agent ranking) has an obvious gap with the first tier, and its performance in reasoning tasks even encountered failures in independent horizontal evaluations.

ByteDance Doubao 2.1-Pro, No.1 in user scale, lacking international verification. It ranks third in SuperCLUE comprehensive Chinese evaluation, with the fastest response speed across all models (the first token delay is about 380 milliseconds), backed by the largest domestic AI user base on the Doubao App. ByteDance is currently advancing the preliminary demonstration of a super-large model with 5 trillion parameters. But it is absent from the Arena international blind test, and no record of this model can be found on the AA index. A candidate who does not appear in the international exam cannot participate in the "global coordinate" ranking.

Direct answer to the question

With the three chains of evidence presented, we can directly answer the question "Who is China's strongest large model".

If we take the completeness and strength of third-party verification as the standard, the answer is Moonshot AI's Kimi K3. It is currently the only Chinese model that has scores in all mainstream international evaluation systems, and all its scores rank among the top in the world: 57 points on the AA Intelligence Index ranks 4th globally (this is the highest position in history for any Chinese model and all open-source models), it takes three top domestic spots and one top 3 spot on the four practical Arena lists, and its actual test performance by FireworksAI is close to Claude Fable 5. It is the Chinese champion voted by the global developer community — on Hugging Face, the 2.8T Kimi K3 topped the trending list on the day of its open-source release.

If we take comprehensive Chinese capability as the standard, the answer is Alibaba's Qwen3.8-Max. It ranks first in total score on SuperCLUE, and is the highest-ranked domestic model on Arena's overall ranking. Its comprehensive performance in Chinese scenarios currently has no rivals, plus the Qwen family's world-leading open-source ecosystem, making it the invisible No.1 from the perspective of industry penetration. But one judgment needs to be reserved: its international standardized report card is still blank, and the official claim of "second only to Claude Fable 5" is currently a "manufacturer assertion" waiting for third-party verification.

So the most honest statement is that the first tier of China's large model industry forms a two-strong pattern with Kimi K3 and Qwen3.8-Max: