HomeArticle

Seed has its own sense of mission

新眸2026-08-14 10:22
The mission of a seed is to learn how to grow in the soil before the season that belongs to it arrives.

Earlier this year, Seedance 2.0 was released without extensive pre-promotion. Within hours, videos generated with it went viral across social platforms at home and abroad: cinematic-level camera movement, native audio-video synchronization, and multi-camera narrative. Even Elon Musk reposted a sample clip on X, leaving a comment: "It's happening fast."

Prior to that, a man named Wu Yonghui had already become the focus of the industry. He earned his bachelor's degree from the Department of Computer Science of Nanjing University, class of 2001, spent 17 years at Google, was a Google Fellow, and served as one of the overall technical leads for the Gemini large model application. A year ago, he left Mountain View, returned to Beijing, and took over Seed, ByteDance's large model research team.

The situation Wu Yonghui faced when he took over was described by media outlets as follows: a research team of thousands of people that had invested tens of billions of yuan for two years to finally develop a foundational model ranked in China's first tier, only to be quickly overtaken by a model developed by DeepSeek, which had a team of just a few hundred people and used far fewer resources. The department head admitted mistakes, and the company's CEO pointed out the issue at an all-hands meeting, noting that the team could have done better.

A year and a half later, the daily average token call volume of the Doubao large model surged to 180 trillion, representing a more than 1500-fold increase over two years. Volcano Engine ranks first in China's public cloud MaaS market with a 49.5% share. Seedance 2.0 has become the global benchmark for video generation models, with gross margins estimated by multiple parties to be over 70%, contributing more than half of Volcano Engine's MaaS revenue.

Doubao 2.1 Pro, released in June this year, matched Claude Opus 4.7 in programming benchmarks such as Terminal Bench 2.1, with a comprehensive usage cost less than 20% of its competitor's. In August, SeedRealtime, a native audio and video full-duplex large model, was released and fully launched on the Doubao App the same day, allowing 380 million users to make video calls with AI that let them watch, listen and speak at the same time overnight.

Behind these numbers lies a quiet transformation of the organization's genetic makeup. An internet company that rose to prominence relying on recommendation algorithms, A/B testing and a horse-race mechanism has spent three years learning to do something it was not good at in the past: basic research. This process has involved misjudgments, personnel turbulence, route disputes, and moments driven by the external environment.

The story of Seed is not only about technology and people, but also about how a company redefines its boundaries during a paradigm shift.

01

The Old Map and the New War

ByteDance did not start working on AI large models particularly early.

After the release of GPT-4 in 2023, the company set up a special large model task force led by Li Hang, then head of the AI Lab. In August of the same year, the Skylark large model completed filing, five months later than Baidu's Ernie and three months later than Alibaba's Tongyi Qianwen. Zhu Wenjia was transferred from the search business line to lead the team that formed the predecessor of Seed.

This tech executive, who was evaluated by Yang Zhenyuan, ByteDance's head of recommendation algorithms, as one of the "top 3 algorithm engineers at Toutiao", has a resume with typical ByteDance characteristics:

He was a former chief architect at Baidu's search department, joined ByteDance in 2015 to work on search, took over Toutiao in 2019, reported directly to Zhang Yiming three months later, pushed Toutiao's DAU to 140 million during the 2020 Spring Festival, and was transferred to Singapore in 2021 to support TikTok's technology development.

He understands large-scale systems, recommendation algorithms, and how to turn technology into products for hundreds of millions of users.

2024 was the year of ByteDance's AI application factory. In May, the Doubao large model family was officially released, and Volcano Engine cut the API price to 0.0008 yuan per thousand tokens, 99.3% lower than the industry average, firing the first shot of the large model price war.

Relying on ByteDance's mature product methodology, traffic system and growth hacking methods, the Doubao App grew from scratch to become the AI application with the highest DAU in China within a year, with monthly active users approaching 60 million in November 2024, leaving the second-place app nearly 50 million behind.

ByteDance is extremely familiar with this set of strategies. Over the past decade, from Toutiao to Douyin, from Xigua Video to CapCut, the same logic has been repeatedly verified: rapid iteration, data-driven decision-making, A/B testing, horse-race mechanism, and saturated investment. Whenever it enters a content or tool track, ByteDance can always overtake others with more refined operations and more aggressive resource investment.

Last year, DeepSeek released R1, and the capability curve of reasoning models saw a steep jump that winter, but ByteDance did not even have a corresponding reasoning model product. The user advantage accumulated at the application layer appeared fragile in the face of generational gaps in foundational capabilities.

As a result, Liang Rubo delivered his second reflection at the company's all-hands meeting that year. The last time he said "the company has become sluggish" and neglected the language models centered on Transformer. This time he admitted that the follow-up to OpenAI o1 was too slow, and the team at the time thought it made little difference whether they moved one month earlier or later.

He said a sentence that was later widely quoted: It is not enough to be a technology company; we have to be an innovative technology company. We must not only apply new technologies well, but also be able to explore and invent new technologies.

This was a crucial cognitive turning point. The success of the internet era was largely built on engineering catch-up: identify the right direction, and use stronger engineering capabilities and resource investment to quickly narrow the gap. But in the large model era, the feedback cycle of the recommendation system is measured in hours, while the feedback cycle of model capabilities is measured in months or even quarters. The former can advance quickly through small steps with A/B testing, while if the architecture choice and training route of the latter are wrong, months of investment will be wasted.

That year, Liang Rubo set a new goal for Seed: pursue the upper limit of intelligence. Take intelligence itself as the most important goal, rather than the DAU of a certain product. This adjustment of goal priority is more fundamental than any change to the organizational structure.

02

The Successor and the Implantation of Research Culture

Last February, Wu Yonghui officially joined ByteDance. He was another external executive who directly joined at the CEO-1 level after Gao Zhun, ByteDance's CFO.

He earned his bachelor's degree at Nanjing University, his doctorate at UCR, and joined Google in 2008 where he worked for more than a decade. In the search ranking team, he developed the work habit of "building large-scale machine learning systems to serve hundreds of millions of users". In 2015, he transferred to Google Brain and was a core author of the GNMT neural machine translation system. He was also a core contributor to later technologies including the Conformer speech architecture, GLaM sparse expert model, and CoCa image-text model.

After Google Brain merged with DeepMind, Wu Yonghui was promoted to Vice President of Research and Google Fellow. In the Gemini team, he was one of the overall technical leads on the application side, leading the implementation of the million-token context window. Google Scholar shows that his papers have been cited more than 73,000 times in total.

DeepSeek-R1 was released one month before he joined ByteDance. The shockwave had not yet dissipated inside ByteDance. Right after Liang Rubo finished his reflection at the all-hands meeting, Wu Yonghui stepped onto the front stage.

In fact, ByteDance's AI organizational adjustment had already started before Wu Yonghui joined. In January 2025, Seed Edge was officially launched, targeting long-term research of more than five years, with the assessment cycle extended to three years. Teams that deliver important achievements can even get "retroactive performance compensation". This is almost an anti-OKR existence in ByteDance, which normally operates on a six-month OKR rhythm.

None of the five directions — next-generation reasoning, next-generation perception (world model), hardware-software integrated model design, next-generation paradigm (beyond backpropagation and Transformer), and next-generation Scaling (Multi-Agent and test-time training) — can produce clear and definite outputs within a year.

After taking office, Wu Yonghui further built a three-layer virtual organization consisting of Edge, Focus and Base: Edge focuses on long-term exploration of more than five years, Focus tackles the core bottlenecks of next-generation models, and Base is responsible for the engineering delivery of current-generation models. Personnel and topics can flow between the three layers: achievements from Edge can be transferred down to lower layers, and long-term topics discovered by Focus can be moved to Edge.

He encouraged communication and exchanges, made data and code libraries fully transparent internally, and broke down the previous barriers where documents between different teams were not visible to each other.

Zhang Yiming also devoted more energy to this matter than many people expected. The founder started to meet one-on-one with authors of important AI papers at the end of 2023, including doctoral students who had not yet graduated. After Wu Yonghui joined, Zhang Yiming returned to the front line to check on technical progress more frequently. He often attended Seed's important technical meetings and occasionally discussed details with researchers.

But the implantation of research culture cannot be achieved overnight. When Wu Yonghui took over and trained the first large model, Doubao 2.0, he gave the entire team a real lesson. This was the largest model Seed had ever trained, with one trillion parameters and a Gemini-like multimodal architecture.

When the parameter scale increased, training stability problems emerged. In the words of one researcher, "it's like building a house: if the foundation is unstable, problems will easily arise as you add more bricks up". Later, multiple teams cooperated and spent three months patching both the model architecture and training data, finally finishing the model before the 2026 Spring Festival.

Weng Jiayi, the head of OpenAI's RL Infra, once said: Every model team's Infra has bugs. Essentially, model companies compete on the speed of fixing Infra bugs, which determines how many ideas can be verified per unit of time, and ideas can be solved as long as the talent density is high enough.

For ByteDance, personnel reshuffling and mixed staffing started very early. During the formation of Seed, ByteDance intensively recruited talent around the world. Jiang Lu, head of Google's VideoPoet, Zhou Chang, former technical lead of Alibaba's Tongyi Qianwen, and Huang Wenhao, co-founder of 01.AI, joined the team one after another.

In the year Wu Yonghui took office, Zhu Wenjia's reporting line was changed from Liang Rubo to Wu Yonghui. The early dual-leader structure was transformed into a single-core architecture where Wu Yonghui leads basic research and Zhu Wenjia is responsible for model applications. The visual generation line also experienced adjustments: text-to-image Seedream, text-to-video Seedance, and 3D generation Seed3D were all placed under Zhou Chang's management.

Entering 2026, personnel reinforcement is still ongoing. Guo Daya, a core author of DeepSeek R1, joined Seed to lead the Coding and Agent direction, and related R&D resources were unified under his management. In June, the Seed Robotics team was also merged into Zhou Chang's jurisdiction, further centralizing the rights and responsibilities of the multimodal and embodied intelligence line.

A detail worth mentioning is the training decision for Seedance 2.0. At a team dinner at the end of 2025, Zeng Yan, the head of the video model team — a young researcher who joined the company through campus recruitment in 2021 — proposed to directly train a 200B parameter model.

There were disagreements within the team: some people thought it was more prudent to train a 100B model first, and training resources were also tight. After discussing with Zhou Chang, Wu Yonghui chose to support her judgment. This seemingly radical choice at the time later proved to be one of the key factors that allowed Seedance 2.0 to open up a generational gap over its competitors.

03

Seed That Does Not Use Distillation

At Seed's all-hands meeting some time ago, Zhang Yiming rarely spoke for a very long time. One core message was: Seed will not use distillation to catch up. Neither closed-source models nor open-source models can be distilled, and the team can accept falling behind temporarily.

According to subsequent reconstructions by multiple media outlets, this decision went through three route disputes inside Seed.

The first time was after the release of DeepSeek-R1 in January 2025, when some researchers proposed to use generated data from leading models to complement reasoning capabilities. The second time was when NVIDIA's Blackwell GPUs were deployed on a large scale at the end of 2025, and the computing power gap between China and the United States was further widened by the new generation of hardware. The voice of "distill if we really can't make it" emerged again.

The third time was after the release of Kimi K3 in July 2026. Moonshot AI pushed the open-source model into the global first echelon with a much smaller team size, and the wavering sentiment inside Seed reached its peak. There was even a compromise proposal that "we can at least distill open-source models".

Zhang Yiming vetoed all three proposals. Judging from the timeline, ByteDance's vigilance against distillation is not a short-term attitude. Back in April 2023, the team introduced GPT API call specification checks, explicitly prohibiting GPT-generated data from being added to the training set, and also conducted internal spot checks to rule out such situations.

Zhang Yiming's exact words at the meeting were "to achieve long-term goals, we should be willing to sacrifice some short-term interests". The underlying logic of this judgment is not complicated. Distillation is a sufficiently effective shortcut that allows models to close the gap on benchmarks within a few months. But it cannot tell researchers how these capabilities are generated, nor can it guarantee where the distilling party should go next when the existing SOTA hits its ceiling.

A model can reach the cutting edge through distillation, but a laboratory cannot become a cutting-edge laboratory through distillation.

Accepting temporary lag means delivering sufficiently strong achievements in other fields to maintain confidence. Seedance 2.0 has played this role.

Video generation is the track where ByteDance is closest to overseas players in terms of starting point. After Kuaishou's Keling adopted the DiT native video architecture, it took the lead for nearly a year in 2024. In the early stage, ByteDance's internal PixelDance took the route of extending 2D UNet to 3D, which later turned out to have a low upper limit. Zeng Yan led the video team from AI Lab to merge into Seed, the technical route was switched from UNet to DiT, and the team and direction converged at the same time.

The success of Seedance 2.0 was attributed by many practitioners to "a victory of data". ByteDance supported this model with a data evaluation team of thousands of people, while many startups in the video generation track only have evaluation teams of dozens of people.

The algorithm team has dedicated personnel to connect with the data team, putting forward clear requirements like data product managers, and algorithm personnel also participate in data construction and cleaning. However, the training data hardly uses content from Douyin, but purchases a large amount of film and television-level materials, which are disassembled into scripts and storyboards using language models, and then targeted training is carried out for scenarios such as motion, indoor spaces, and game frames.

Commercial returns came quickly. During the 2026 Spring Festival, Seedance 2.0 was fully integrated into Doubao, CapCut and Jimeng. The C-side waiting queue once lasted as long as ten hours, and Jimeng generated its first batch of revenue through membership priority queuing, with revenue reaching about 140 million yuan in March and rising to 210 million to 220 million yuan in April.

Volcano Engine immediately opened up enterprise APIs. With a pricing of about 1 yuan per second for 720P, Seedance is indeed more expensive than other domestic video models of the same generation. But quality speaks for itself: cases where leading short drama companies such as Chinese Online and Jiuzhou Culture recharge 50 million yuan at one time are not rare.

Tan Dai confirmed in an interview that more than half of Volcano Engine's MaaS revenue in 2026 was contributed by Seedance. The high pricing, high gross margin and relatively loose competitive landscape of video models allow SOTA model advantages to be directly converted into revenue.

More importantly, it has formed a closed loop with ByteDance's original business: models are provided to content companies, which use Seedance to make AI short dramas, while Hongguo and Douyin undertake content supply and support producers to place ads. A flywheel spanning from model capabilities to content production and then to advertising revenue has started to operate first in the video modality.

At the Volcano Engine FORCE conference