In a rare public statement, Zhang Yiming fully escalates the AI battle.
Zhang Yiming, who has long remained low-profile, has once again become the focus of China's AI circle.
On August 5, The Information reported that at the all-hands meeting of Seed held last month, ByteDance founder Zhang Yiming set a strict rule for this core team responsible for large model R&D: They must not rely on distilling the cutting-edge models of other companies to improve their own capabilities.
At this moment, the competition of large models in China is in full swing. The rankings on the leaderboards keep changing, and as soon as a new model is released, all companies immediately conduct research, catch up and iterate.
For a team that is still in the catch-up phase, distillation is one of the fastest ways to narrow the gap. However, Zhang Yiming voluntarily blocked this path.
01
Prefer to fall behind temporarily
Since the second half of 2024, Zhang Yiming has participated in the review and discussion meetings of the Seed core technical team once a month.
The Information reported that now he devotes roughly half of his working time and energy to Seed. But Zhang Yiming rarely talks about AI in public, even in internal meetings, he mostly just listens and raises follow-up questions, and almost never takes the initiative to express his stance.
It was not until July 2026, when the team discussed again whether to catch up with competitors through distillation, that he rarely spoke up.
According to employees attending the meeting, Zhang Yiming made it clear that even if Seed falls behind temporarily, it cannot rely on distilling the cutting-edge models of other companies to improve its capabilities.
How tempting is distillation? To put it simply, it means finding a stronger model as a teacher, and letting your own model learn the answers and problem-solving methods it provides.
This technology is not new. As early as 2015, "the father of deep learning" Geoffrey Hinton and others systematically promoted knowledge distillation.
In the past, it was mainly used to compress models and reduce costs. After the emergence of ChatGPT, things have changed.
When the most advanced large models have become the core technical assets of various companies, distillation has also been involved in intellectual property rights, technological originality, and Sino-US AI competition.
But the question debated inside Seed is more straightforward: since others have already worked out the answers, why not learn from them first?
Especially in the field of large language models, ByteDance is under considerable pressure. Whenever a stronger open-source model appears, discussions about distillation will heat up again.
In the end, Zhang Yiming made the final call: To achieve long-term goals, we should be willing to sacrifice some short-term interests.
The Information attributed this decision in part to the risks TikTok is facing in the United States. But on August 6, ByteDance CEO Liang Rubo responded to this statement at ByteDance's mid-year all-hands meeting.
He admitted that ByteDance's large language model is indeed lagging behind at present, and also acknowledged that the gap with leading overseas models is still widening. But for the explanation that "external pressure forced ByteDance to adhere to independent R&D", Liang Rubo directly used two words: "Pure speculation."
According to his explanation, the real reason is that he hopes the team can learn delayed gratification and lay a solid foundation of basic skills.
What Zhang Yiming blocked is not only the path of American models. According to domestic media reports, ByteDance also has strict internal restrictions on the use of other open-source models for distillation, and even strengthens control through API detection and other methods.
This speaks volumes. What he is really worried about is that Seed will get used to walking on the road paved by others. Distillation can help a team quickly approach the existing goals, but it can hardly tell it where the next goal lies.
Learn from Claude today, catch up with GPT tomorrow, and chase a new model the day after tomorrow. The ranking may get better and better, but the research direction, training method and even evaluation criteria are all following others' lead.
Over time, ByteDance can become a very strong pursuer, but it will hardly become the first person to raise new questions. Therefore, Zhang Yiming would rather slow down a little for the time being.
But closing the shortcut is only the first step. If you really want Seed to find its own way, he has to get rid of some of the most familiar successful methods of ByteDance in the past.
02
Rebuild ByteDance
To understand Zhang Yiming's logic, we have to go back to 2012.
That year, Toutiao was just starting. One night at the end of the year, in the office on the 6th floor of Jinqiu Jiayuan, Zhichun Road, Beijing, Zhang Yiming gathered all product managers and R&D personnel for a meeting.
The core of the discussion was only one thing: since we want to build an information platform, should we really build a personalized recommendation engine?
At that time, not many Chinese startups were determined to build a recommendation engine, and Zhang Yiming and his team could not do it either. Some people worried: "We don't have the genes and capabilities for recommendation."
Zhang Yiming's answer was very simple: "We don't know recommendation, but we can learn it."
Later, ByteDance really started to learn from scratch. The machine records what users clicked, how long they stayed, what they like, and then constantly guesses what this person wants to see next.
Toutiao changed news distribution with this set of algorithms, and Douyin brought it into the short video era.
Many years later, when Zhang Yiming recalled the early days of starting his business, he said a classic sentence: "Looking back, many of our methods were not good at the beginning, but we worked very hard, were very focused, and great effort brings wonders."
"Great effort brings wonders" later almost became a trait of ByteDance in its early days. Launch the product first, let data speak, quickly test errors, and iterate continuously. Once the correct direction is found, resources will be mobilized immediately to amplify the advantages rapidly.
This is true for both Toutiao and Douyin. In the mobile Internet era, this set of methods allowed ByteDance to catch up from behind again and again.
But by 2021, Zhang Yiming suddenly admitted that he was a little unable to keep up. In May of that year, at the age of 38, he announced that he would step down as CEO of ByteDance.
In the internal letter, Zhang Yiming did not shy away from saying that he had been "living on past gains" to a large extent in the past few years. Before 2017, he was still able to continuously follow up on new developments in machine learning technology. In the three years after that, he rarely had time to study systematically, and even found it difficult to keep up in some technical seminars.
It is very rare for a person who started a business with algorithms to publicly admit that he cannot keep up with the latest algorithms.
Zhang Yiming decided to hand over the daily operation of the company to Liang Rubo, so that he could free up more time to study, think and do research. He set a very long time frame for himself: Ten years.
At that time, from the outside world's perspective, this seemed like the exit of a founder. Five years later, it looks more like a role swap.
Liang Rubo took over the huge daily operation of ByteDance, while Zhang Yiming paid more and more attention to AI and cutting-edge technologies.
In the past two years, he has traveled frequently between Beijing and Singapore. In Singapore, he communicates with AI researchers; back in China, he participates in discussions of the Seed technical team, watches model training, and keeps an eye on long-term research directions.
This time, the problem he is facing is completely different from that in 2012.
The recommendation algorithm in 2012 had already been led by others. If you don't know it, you can learn it. But at the cutting edge of today's large models, there are still a large number of problems for which no ready-made answers exist.
What is the architecture of the next-generation model? How should the world model develop? Where can Agents go? To answer these questions, relying solely on "great effort brings wonders" is no longer enough.
ByteDance began to loosen restrictions for some people.
In early 2025, Wu Yonghui, former Vice President of Research at Google DeepMind, joined ByteDance to be in charge of Seed basic research. During the same period, ByteDance launched the Seed Edge program, which focuses on AGI topics with longer research cycles and more uncertain results.
▲ Dr. Wu Yonghui
This team canceled quarterly OKRs and semi-annual assessments, and also received independent computing power support. For a company that has long emphasized quarterly goals, rapid iteration and data feedback, such an arrangement is very rare.
Wu Yonghui is also changing the daily operation of Seed.
According to LatePost reports, there used to be information barriers between different research groups. If one team wanted to view the documents and codes of another team, it sometimes had to go through multi-layer approval. After taking office, Wu Yonghui began to promote information sharing, and also encouraged researchers to publish papers and blogs to actively display research results to the outside world.
At one all-hands meeting, he even reminded everyone to "decorate" their personal homepages.
Changes soon emerged. According to statistics, in the three months after Wu Yonghui joined, Seed published more papers than in the whole of 2024.
A more vivid change took place in the canteen. Some Seed researchers recalled that they often saw Wu Yonghui carrying a dinner plate, sitting directly opposite young researchers, and asking about the progress of their projects while eating.
These small changes are adding something to Seed that ByteDance rarely had in the past: Patience.
Allow a project to have no results in the short term, allow some research to be temporarily uncommercialized, and allow researchers to spend time solving problems that may be truly important in a few years.
Therefore, Zhang Yiming's rejection of distillation from external models and the emergence of Seed Edge are actually two sides of the same thing. On the one hand, he closed the ready-made answers, and on the other hand, he left time and computing power for those who are looking for new answers.
03
Where is the AI war going?
On July 31, ByteDance released Seedance 2.5.
It has not been long since Zhang Yiming set the tone of "not relying on distillation of external models" at the Seed all-hands meeting.
This model is also very suitable for observing where ByteDance is at present. Compared with Seedance 2.0 released earlier this year, the new model can generate 30-second videos at a time, and can refer to up to 50 image, video and audio materials at the same time.
But what is more interesting is that it has begun to step out of the short video field.
XCMG Group uses Seedance to generate industrial operation training and process SOP videos. XPeng Motors integrates it into its internal AI design platform. Diffy Aero uses it to "build training grounds" for flying robots, simulating complex environments such as obstacle avoidance and indoor autonomous pathfinding.
▲ Winning work of XPeng's first AI Open Competition, the future concept smart cockpit generated by the team based on Seedance 2.0
An AI model that used to be mainly used to generate cinematic images has begun to work for excavators, cars and robots.
The video model is the battle where ByteDance has performed best so far.
This is not surprising. Douyin, TikTok and Jianying have allowed ByteDance to accumulate more than ten years of experience in the video field. ByteDance is very familiar with what kind of pictures look good, how users edit videos, and what tools creators lack. In the era of generative AI, these product, data and engineering experiences have begun to translate into model advantages.
Another field where ByteDance has achieved obvious advantages is Doubao.
QuestMobile data shows that as of March 2026, the monthly active users of Doubao reached 345 million, that of Qianwen was 166 million, and that of DeepSeek was 127 million. The number of Doubao's monthly active users has exceeded the sum of the latter two.
What is more exaggerated is the call volume.
In May 2024, the average daily Token call volume of the Doubao large model was only 120 billion. By April 2026, this number exceeded 120 trillion; in June, it further rose to 180 trillion.
In more than two years, it has increased by 1500 times.
This is the most familiar battlefield for ByteDance. Once the model is available, it will be quickly pushed to a huge user group, and then real usage will generate more feedback.
In June this year, Doubao officially launched paid subscription, with monthly fees ranging from 68 yuan to 500 yuan. After having hundreds of millions of users, ByteDance began to answer another question: How on earth can AI make profits?
However, what Zhang Yiming is really anxious about lies at a more fundamental level.
At ByteDance's mid-year meeting on August 6, Liang Rubo presented a very contrasting report card for the AI business: Doubao remains competitive, Seedance continues to lead, but the gap between the large language model and the leading overseas models has widened.
ByteDance chose to continue to increase investment.
On August 7, Financial Times revealed that ByteDance is pre-training a new model with a maximum parameter scale of up to 10 trillion, which is more than three times that of Kimi K3.
The project is still in the early stage, and the final training scale and capability ceiling are not yet finalized. But the move is already big enough. Zhang Yiming just closed a shortcut, and ByteDance turned around and continued to bet on the more capital-intensive frontal battlefield.
There are still two tough nuts to crack: Coding and the world model.
Coding, that is, the ability of AI to write code, directly affects the upper limit of Agents to complete complex tasks. 36Kr reports that ByteDance's investment in this direction in 2026 is second only to the world model, but it has not yet formed a leading advantage similar to Seedance.
In the past, due to the limited capability of its own Coding model, many of ByteDance's business teams were unwilling to use Seed-Code. In the early stage, its AI programming tool Trae also accessed external models such as DeepSeek and Claude.
This brings another problem: insufficient usage makes it difficult for real development data to flow back, and it is more difficult for the model to continuously improve through actual scenarios.
Since 2026, ByteDance has begun to promote more internal businesses to use the Seed model, hoping to re-establish the flywheel of "usage - feedback - training".
The world model is more like a battle for the future.
ByteDance did not set up a research team on a small scale until 2025. In 2026, Wu Yonghui set a goal: to release at least one version of the world model by the end of the year, with performance benchmarked against Google Genie 3.
Investment has increased significantly, but as of the beginning of 2026, internal evaluation showed that the comprehensive performance of ByteDance's world model is still about 10% behind the world's leading level.
At this point, ByteDance's AI battle situation is very clear.
Seedance has broken through, Doubao has gained users, the basic large model and Coding are