Following Seed and Flow, ByteDance has established another first-tier AI division that is directly focused on "data" | 36Kr Exclusive
Intelligent Emergence has learned from multiple independent sources that ByteDance has recently established a new first-level department — AI Data and Security, which is parallel to departments including Seed, Flow, and Douyin, and is led by Wang Yinglei (Adam Wang).
This is another first-level AI-focused department established by ByteDance after the two AI first-level departments Seed and Flow were set up at the end of 2023. Previously, Wang Yinglei, the head of the new department, was the Head of Platform Responsibility and Head of Live at TikTok. "The live streaming business led by Wang Yinglei used to be one of the most important revenue sources of TikTok, and he has made outstanding achievements inside ByteDance," said a person close to ByteDance.
Intelligent Emergence has reached out to ByteDance for verification on the above news, and there is no response as of press time.
After models and products, ByteDance is targeting AI data directly this time.
According to the understanding of Intelligent Emergence, one of the predecessors of this new department is the Global Data team, a data team established in 2023 by Fu Yue (alias Yuyi), a member of TikTok's founding team, which originally had a scale of about 100 people. This lean team initially served international businesses including Dola (the overseas version of Doubao) and TikTok, and later took charge of data procurement and quality control for Seed's model training. The team has multiple internal functions including product manager, data engineering, procurement, quality inspection operation, security and compliance.
In addition to Global Data, the AI Data and Security department has also integrated multiple previously scattered AI data departments, including the Group Data Middle Platform DMC, and team personnel of the AI data platform AIDP under Flow.
Multiple insiders told us that the integration and official establishment of this department started in early June 2026. Since it involves the integration of multiple departments whose businesses largely overlap with each other, ByteDance is still sorting out the structure and optimizing personnel continuously.
Several people close to the department told us: The integrated new department will be a huge data team spanning from the foundational model to business teams, whose core function is to provide cross-modal data services for all large models of ByteDance. The responsibilities of the team also cover the entire data production process: standard formulation, sourcing and procurement, synthesis and cleaning, quality evaluation, etc.
At the all-hands meeting of Seed at the end of July, Zhang Yiming, who had not appeared in public for a long time, stated that ByteDance would "resolutely not adopt distillation" in the tough work of large models. At the ByteDance all-hands meeting on August 5, CEO Liang Rubo also said, "ByteDance's large language models will adhere to independent R&D, do a solid job in basic skills, accept short-term backwardness, insist on long-term optimization, and the most important thing is not to deviate from our established direction." These decisions will inevitably lead to the continuously rising importance of data in ByteDance's large model research.
Looking beyond ByteDance, we are also seeing structural changes taking place in the global AI industry: when public internet data is exhausted, the core variable of model competition is shifting from algorithms and computing power to high-quality data in the real world.
ByteDance's Massive Data Empire
ByteDance's investment in AI data has always been the most resolute among major domestic tech companies.
One important reason is that Seed set the principle of "no distillation" from the very beginning of its establishment. ByteDance's goal for almost all its models is to reach the global first tier or even SOTA, which cannot be achieved through distillation. If all data needs to be synthesized, purchased and cleaned by itself, a huge team is bound to be required to provide support.
How huge is this team? Intelligent Emergence once exclusively reported that inside ByteDance, there are as many as a thousand people in the team that only does model data evaluation for Seedance. Behind every Seedance algorithm engineer, there are often more than ten data colleagues providing support.
In contrast, for many leading startups in the video field, an internal evaluation team of dozens of people is already considered a relatively large investment. The success of Seedance 2.0 is also called "a victory of data" by many practitioners.
This has made ByteDance more determined to invest in data.
The first is the budget aspect. Intelligent Emergence once exclusively reported that for the training of world models and Coding models, ByteDance's data budget in early 2026 has exceeded the tens of millions of dollars level, and "the budget can be increased at any time if it is considered insufficient".
Intelligent Emergence also learned that ByteDance's data team has now adopted a horse race mechanism, and the internal team is divided by direction — world model, coding, advanced difficult disciplines, etc. Under the background that the commercial logic of models is becoming increasingly clear, each data project of ByteDance also needs to calculate ROI.
Apart from ByteDance, other major tech companies are strengthening their data investment. Tencent, which has made real efforts in large models since last year, has frequently poached people from ByteDance's data team with salaries as high as three times in the past six months.
Both Alibaba and Tencent have seen a significant increase in their budgets for data procurement, and are continuously increasing their investment. A person in the data industry told Intelligent Emergence that major tech companies currently have certain data exclusive strategies, such as setting an exclusive period for data sets, or buying out core personnel of suppliers in stages.
"Data is the most decisive variable in the current competition of large models." This has become a consensus in almost all star sub-sectors such as large language models, embodied intelligence, world models, and AI4S.
However, the current situation of the data market in 2026 is that AI companies have strong demand, and they are facing the bottleneck of scarce high-quality data supply from pre-training to post-training.
For ByteDance, the next challenge is how to make a data organization of a thousand people still able to respond quickly to the cutting-edge changes of models. After all, at the stage when the capability boundary of models is still changing rapidly, the most scarce data today may lose its value in half a year.
Global Model Competition is Heading for a Data Arms Race
In the past year, the global tough work of large models has faced a trend: on the premise that the reasoning paradigm is converging slowly, the capability improvement brought by algorithm architecture is slowing down. Data has become the most important variable that determines the upper limit of model capability, even the only one.
Compared with Chinese tech giants such as ByteDance, overseas large model companies have invested several times more in data. The external data budget of leading large model companies in Silicon Valley is expanding rapidly at a magnitude of tens of billions of dollars every year. For example, Anthropic allocated a budget of more than 1 billion US dollars for RL data in 2025 alone.
This has made data one of the fastest growing tracks in Silicon Valley in 2026. The most typical example is Mercor, a star company in Silicon Valley, whose annualized revenue rose from 500 million US dollars last year to 2 billion US dollars by the middle of this year, 91% of which came from leading model companies such as OpenAI and Anthropic, and its valuation has soared to 20 billion US dollars. This company was established only three years ago.
Why has data suddenly become so important? The fundamental reason is that the public data on the Internet has almost been exhausted.
In the past three years, the data required for model training has undergone a transformation from the public domain to the private domain. There are a large number of reports and documents on the public Internet, but the process of how humans produce these results is missing, such as how to understand vague intentions, find context, make mistakes and make corrections.
For example, in addition to Coding, there are high-end and sophisticated fields tasks for doctors, lawyers, scientific research and other scenarios, which general models have not been solved very well. When a model wants to improve its capability and work like a real expert, it also needs a large amount of naturally generated "process data". These data are often hidden in the real work processes, such as the code base deposited inside enterprises, the traces left by employees in their daily operations, and so on.
The data required by AI is shifting from the public domain to the private domain. The collection of a large amount of private domain data has become a new point of contention for large model companies.