Zhang Yiming retreats in order to advance
On August 6, Liang Rubo, CEO of ByteDance, admitted at the company's all-hands meeting that the gap between Seed's large language model and leading overseas models is widening. He asked the team to accept being behind for a period of time, adhere to independent research and development and long-term optimization. Zhang Yiming explicitly opposed distilling competitors' models at an earlier internal meeting.
On the same day, LatePost disclosed that ByteDance is discussing training a large model with a parameter scale exceeding 5 trillion. As a reference, Alibaba's Qwen 3.8-Max and Moonshot AI's Kimi K3 listed in the report have parameter scales of 2.4 trillion and 2.8 trillion respectively. This plan is still in the early discussion stage and may not be released in the end.
These are two seemingly contradictory signals, but in fact the former resets the catching-up timeline, while the latter determines the catching-up method: ByteDance is not prepared to quickly fill the rankings with the output of competitors' models, but tries to concentrate computing power, data and organizational resources on a larger-scale pre-training.
This is a route with higher cost and more unpredictable results. The parameter scale only indicates the investment direction, which does not equal the final capability of the model. What ByteDance really wants to verify is whether it can turn resource advantages into a training system and organizational capability that independently generate cutting-edge capabilities.
The bet of "jumping to several times that of peers"
According to reports, the over 5 trillion-parameter model will be led by Xiang Liang, head of Seed Foundation, and cooperate with Shen Ke, head of large language model pre-training data. Both of them come from ByteDance's search, advertising and recommendation system. To this end, Seed is re-dividing responsibilities and allocating resources.
The direct reason for promoting this plan is that Seed's performance in the first half of the year did not meet internal expectations. Seed 2.0 released in mid-February received limited market response, while domestic models such as GLM-5 and Kimi K3 have made significant progress in programming, tool invocation and complex tasks.
Instead of pinning its hopes on winning while catching up, ByteDance might as well move to a bigger gambling table.
The bet is first reflected in the engineering complexity. For models adopting the Mixture of Experts architecture, total parameters, single activation parameters, training computation and inference cost are different indicators. 5 trillion parameters can expand model capacity, but cannot automatically solve the problems of data quality, training stability and capability emergence. If the parameter scale is not upgraded simultaneously with the architecture, data and training methods, you may only get a more expensive model in the end.
Therefore, 5 trillion is more like ByteDance's statement on resource allocation mode, rather than a technical check cashed in advance. The "search, advertising and recommendation" background of Xiang Liang and Shen Ke also shows that this task is not only about algorithm research, but also includes large-scale training systems, data engineering and cross-team collaboration. The larger the model, the inefficiency of any link will be amplified.
Being behind has begun to affect revenue
Another reason why ByteDance is eager to turn the tables lies in its books.
First look at the gap on the technical side. In mid-February this year, Seed 2.0, the key model launched after Wu Yonghui took charge of Seed, was officially released, with limited market response. Almost at the same time, GLM-5 open-sourced by Zhipu was regarded by the market as the first domestic model that can rival Anthropic Opus series. Multiple third-party evaluations believe that Kimi K3 released in July is close to overseas closed-source flagship models.
The most painful part of the gap is Coding. During the Spring Festival, Anthropic quickly opened up the market of programmers and B-end customers with the programming capability of Opus series. According to public reports, its ARR soon approached and surpassed OpenAI. Zhipu and Moonshot AI, with their continuously improved Coding capability, saw their ARR exceed 1 billion USD and 300 million USD respectively. Chinese big tech companies realized at this point that they had missed the Coding window period.
Then look at ByteDance's revenue structure. Volcano Engine is currently the largest model API seller in China, accounting for about half of the market share, with revenue of about 150 billion yuan in 2025, and its internal target for this year exceeds 400 billion yuan. But according to LatePost's report, among the token consumption of Doubao large model, more than half comes from the two multimodal generation models Seedance and Seedream, and the proportion of language models is not high. As the growth of the short drama industry peaks, the token consumption and revenue growth rate of Seedance slow down. The token consumption of Doubao large model was 120 trillion in March and 180 trillion in June, which is lower than the original target of 250 trillion to 300 trillion.
In short, ByteDance's AI revenue is skewed. Multimodal generation earns money from content consumption, whose demand follows the industry cycles of short dramas, e-commerce materials and others. Language models, especially Coding, earn money from productivity, with customers being developers and enterprises, who have much higher stickiness and unit price. The revenue target of 400 billion yuan cannot be achieved by video generation alone, and the short board of language models must be filled.
No distillation, leaving the cost of catching up to itself
Two weeks ago, Zhang Yiming and Wu Yonghui, head of Seed, attended the Seed all-hands meeting together. According to media reports, Zhang Yiming said at the meeting that training large models is inherently difficult, and the team can accept being behind for a period of time. He recognized that Coding is the current key direction, but also reminded the team that programming is only one of the current hot spots, and all research resources should not be driven by a single scenario.
His attitude towards distillation is clearer: do not rely on the output of competitors' models to improve Seed.
Two practices need to be distinguished here. Distillation and synthetic data are mature model training technologies, and top labs generally use their own models to generate data to train subsequent versions. What ByteDance opposes is to use closed-source models such as Claude or open-weight models such as Kimi K3 as "teachers" to obtain outputs in batches, and then use these data to make up for its own capabilities of reasoning, programming and tool invocation.
This principle was not formed temporarily. As early as April 2023, ByteDance required that data generated by GPT should not be added to its own model training set, and checked for unauthorized use through API call inspection and output similarity sampling. Since then, Seed has discussed many times whether to use external model distillation, and all were rejected in the end.
Refusing to distill competitors means that ByteDance has given up a faster catching-up path. Discussing the 5 trillion parameter means that it is willing to use more computing power, data and time to bear this cost. The two decisions actually share the same goal: Instead of approaching along others' capability curves, ByteDance wants to jump to the next capability platform to wait for competitors.
ByteDance has already applied this strategy successfully on video models. In 2025, the mainstream view in the video generation field was that there was little room for improvement by continuing to expand pre-training along the DiT architecture, and most teams shifted their focus to post-training. Kuaishou's Lingke also tried to train a larger model, but turned back halfway. Seedance did the opposite: it pushed pre-training to the extreme, built the first video generation model that fully adopted the MoE architecture with 200 billion parameters. After its launch in February 2026, it was recognized as the world's most high-performance video model, and then became the revenue base of Volcano Engine's MaaS business.
But this case cannot directly prove that the 5 trillion-parameter language model will also succeed. At that time, Seedance bet on a direction where consensus receded, with fewer competitors and only more than ten people in the core algorithm team. Expanding language models is a route that OpenAI, Anthropic, xAI and top Chinese labs are all advancing. ByteDance has no obvious information advantage, and can only bet on whether its resource gap can be converted into training efficiency and model capability.
Great Efforts Bring Miracles 2.0: From Horse Racing to Heavy Betting
To win this battle, ByteDance is changing its most well-known organizational strategy.
ByteDance's well-known strategy of "Great Efforts Bring Miracles" used to rely on horse racing: arranging multiple teams to conduct parallel trial and error in the same direction, and then concentrating resources based on market feedback. ByteDance launched a large number of information applications in the early stage, and later entered the short video field with Volcano Engine, Douyin and Xigua Video at the same time, all of which are representatives of this mechanism. The essence of the horse racing mechanism is to use the number of teams in exchange for success rate, but the premise is low trial and error cost and fast market feedback.
Cutting-edge large models do not follow this rule. The cost of a single training is measured in hundreds of millions of yuan, and the feedback cycle is measured in years. If ten teams work on ten directions, each direction will have insufficient investment.
Also according to media reports, ByteDance's senior management is pushing Seed to cancel internal horse racing in the same direction, gather R&D resources, clarify team responsibilities and break down departmental barriers. In order to improve Coding capability, Zhang Yiming personally invited Guo Daya, a core researcher of DeepSeek, to join Seed to take charge of special training. Relevant R&D resources have also been uniformly allocated to him. It is under this strategy that relevant resources of Volcano Engine, Lark and Doubao are integrated.
This set of organizational adjustments solves the problem of scattered investment, but cannot eliminate research uncertainty. After the reduction of horse racing, a single route obtains more resources, and the cost of wrong direction judgment will also be greater. For ByteDance, the 5 trillion-parameter model is not only a technical training, but also an organizational training: It requires the product company that used to be good at rapid trial and error to adapt to basic research with longer cycle and slower feedback.
Zhang Yiming said that ByteDance "should be willing to sacrifice part of short-term revenue for long-term goals". The real test of this sentence is not on the model ranking list, but on the income statement.
Liang Rubo said at the all-hands meeting: The large model has great market potential, so there is no need to worry about economic returns, "but revenue is still important". With the annual target of 400 billion yuan, lower-than-expected token growth rate, and multimodal revenue accounting for more than half, it is easy to accept technological backwardness, but difficult to accept the revenue gap brought by backwardness.
At present, what can be confirmed is only that ByteDance has chosen a more expensive catching-up method. Whether 5 trillion parameters can bring a new capability platform, the answer will be given jointly by the model, customers and revenue.
This article is from the WeChat official account "Emphasize Next" (ID: leo89203898), author: Qing Yun, editor: Xiao Bai, published with authorization from 36Kr.