Over 7 trillion, DeepSeek V4 Flash ranks first globally in weekly call volume.
The weekly global AI large model call volume ranking (July 27 to August 2) released by OpenRouter, a multi-model aggregation platform, shows that DeepSeek V4 Flash ranks first with a call volume of 7.22 trillion Tokens.
According to a post from the open source project team OpenCode, the call volume of DeepSeek V4 Flash on its platform has surged sharply. On August 1 alone, the model processed 8 trillion Tokens in a single day, of which 5 trillion was consumed through the free trial quota, and 3 trillion was paid for by developers through the OpenCode platform.
OpenRouter data shows that the total global call volume of AI large models last week was 56.8 trillion Tokens, a month-on-month decrease of 2.07%. Among the listed AI large models, the weekly call volume of Chinese AI large models reached 28.13 trillion Tokens, a month-on-month decrease of 14.76%. In the same period, the weekly call volume of US AI large models was 4.38 trillion Tokens, a month-on-month increase of 87.18%. The weekly call volume of Chinese large models has surpassed that of the US for 14 consecutive weeks, firmly ranking first in the world.
On August 5, a query of OpenRouter's this week's ranking by Jiemian News found that as of press time, DeepSeek V4 Flash 0423 still ranks first with a call volume of 6.92 trillion Tokens. The model was first launched in April this year, focusing on balancing inference speed, long-context capability and call cost, and mainly provides high-throughput inference services for developers and commercial enterprises.
OpenRouter's this week's call volume ranking (as of press time)
Ranking second is MiMo-V2.5 launched by Xiaomi, with a weekly call volume of 5.1 trillion Tokens, a 52% month-on-month decrease from the previous week. Its traffic has seen a significant pullback, and market popularity has declined to some extent.
This Xiaomi model officially entered public beta on April 23 this year, and completed full-series open source at the end of April. It adopts the Mixture of Experts (MoE) architecture, with total parameters exceeding the trillion level, 42B activated parameters, and a standard 1 million Token ultra-long context window, which also covers full-modal interaction of text, voice and image. Low inference cost and stable agent operation capability are its core advantages that attract a large number of overseas developers to access.
During the week from July 20 to July 26, MiMo-V2.5 once topped the OpenRouter weekly call volume ranking, with a single-week call volume of 10.5 trillion Tokens, a 12% month-on-month increase.
The third place currently goes to Hunyuan Hy3, a large model self-developed by Tencent, with a call scale of 5.01 trillion Tokens. Hy3 is Tencent's self-developed general large model. The traffic of this model remains flat compared with the previous period, with a growth rate of 0%. Hy3 was once the model with the strongest growth momentum in the ranking of the week ending July 26, when its weekly call volume reached 3.94 trillion Tokens, with a month-on-month increase of over 999%.
Hy3 was officially open sourced on July 6, with total parameters of 295B and only 21B parameters activated per single inference. It adopts a fast-and-slow thinking integrated architecture, supports a maximum of 256K context, and has greatly iterated its capabilities in code generation and intelligent interaction. Tencent Hunyuan once posted that as of July 15, the total call volume of Hy3 has increased by more than 68 times compared with the previous generation model Hy2.
The fourth place also belongs to the newly launched V4 Flash 0731 model from DeepSeek, with a weekly traffic of 3.45 trillion Tokens. The fifth place is GPT-5.6 Luna launched by OpenAI, with a weekly Token call volume of 2.99 trillion, achieving an explosive growth of 738%.
Among the top five seats in the ranking, domestically developed large models take the top four positions, showing extremely high market call popularity. However, the strong catch-up of OpenAI's newly iterated model means that the competition among global leading manufacturers has not slowed down.
Looking at the top ten, DeepSeek occupies three seats. In addition to the two V4 Flash versions, DeepSeek V4 Pro ranks sixth with 2.97 trillion Tokens. Models from Anthropic, Google, NVIDIA, MiniMax, StepFun, Zhipu AI and many other enterprises are all on the list. Free models are occupying an increasing share of traffic. Free models such as NVIDIA Nemotron 3 Ultra, Poolside Laguna S 2.1, and Inclusion AI Ling-3.0-flash have achieved rapid growth, and many free models have a month-on-month increase of over 50%.
It is worth mentioning that the Kimi K3 model, which previously received much attention from Elon Musk, temporarily dropped out of the top ten this week and currently ranks 12th. The model has a total parameter of 2.8 trillion and is equipped with a 1 million Token ultra-long context window, making it the largest open source large model in the world by parameter scale at present.
Different from the past competition mode of relying on low prices to obtain traffic, current domestic models attract overseas developers by relying on continuously iterated inference performance and stable service capabilities. But challenges also exist: OpenAI and Anthropic continue to launch new versions of models, constantly narrowing the performance gap; multiple models of Google's Gemini series stably occupy the top positions in the ranking, and overseas giants still have a solid basic market.
The differentiation within the track is also worthy of vigilance. Some domestic models have seen a sharp decline in call volume in this period, and Xiaomi MiMo-V2.5, MiniMax M3, and Zhipu GLM5.2 have all experienced traffic pullbacks of varying degrees. As the popularity rotation speed continues to accelerate, once the product iteration rhythm slows down, it is very easy to lose developer traffic quickly.
An analyst in the artificial intelligence industry pointed out that the overseas independent developer market has entered white-hot competition, and it is difficult to maintain advantages only by relying on a single version update. How to continue to iterate and efficiently control inference costs will determine whether each major model can retain users for a long time.
This article is from the WeChat official account "Jiemian News", written by Song Jiannan, and published with authorization by 36Kr.