HomeArticle

Why does Zhang Yiming oppose distillation?

嗅态2026-08-09 09:02
As competition for large models enters the deep-water stage, the organizational approach is also facing adjustments.

In July, ByteDance's Seed team held an internal meeting. The topics discussed included the reasons behind the integration of Volcengine, Doubao and Feishu — that is, only by pooling resources can the company gain advantages in computing power and data.

Earlier in August, more critical information was reported by multiple media outlets at home and abroad: Zhang Yiming, founder of ByteDance, talked about "model distillation" at the meeting, and he opposed the practice of model distillation.

In the industry, model distillation is widely regarded as an alternative solution to improve efficiency. By learning from leading models, smaller models can acquire partial capabilities in a shorter time, thus reducing computing power costs and shortening R&D cycles. For latecomers, it is a technical option with quick results and low costs.

Distillation can help a model catch up with an existing target faster, but it can hardly answer what lies beyond that target. The content of the meeting shows that ByteDance has clarified the trade-offs of the Seed team in the R&D path of foundational models.

In the past six months, AI large models around the world have flourished. Among them, ByteDance has continued to make layouts in the multimodal field. The launch of the video generation model Seedance sparked heated discussions in the market, profoundly influencing the AI video generation domain. In the language model field, overseas products including Claude and Gemini have continuously raised the performance ceiling, while domestic players such as DeepSeek, Zhipu AI, and Moonshot AI have also carried out intensive iterations in vertical directions including reasoning and coding.

Facing uncharted territory, Zhang Yiming chose to let the Seed team find the answers on their own. The cost of this choice is more computing power, longer training cycles, a large number of experiments that may yield no results, and organizational patience that is beyond the comprehension of ordinary people. Such a choice may not be immediately reflected in the AI large model rankings, but it determines whether a company has the ability to step into areas where no ready-made answers exist.

Looking around, as China's large model industry develops to the present, the divergence of technical paths among different companies may lie right here. ByteDance's situation is exceptionally different from that of other large model companies.

The Capability Competition Behind Distillation

ByteDance's restrictions on model distillation are more thorough than the outside world imagines.

The restrictions cover the entire external model ecosystem. Cutting-edge closed-source models from the United States cannot be used for distillation, and open-source models are also included in the restricted zone. To implement the rules in the R&D process, ByteDance has also strengthened management through methods such as API call detection, reducing the space for R&D teams to improve their capabilities with the help of external model outputs.

On the surface, Zhang Yiming's opposition to distillation is a trade-off of a technical path. From the perspective of Seed's development stage, it is essentially a choice of the R&D path for large models.

Zhang Yiming's judgment is very clear. Training large models is a high-difficulty, long-cycle project, and model capabilities ultimately need to be built on the long-term accumulation of pre-training, data, algorithms and engineering systems. Even if Seed's current outputs still lag behind the global cutting-edge level, ByteDance is willing to bear the phased backwardness and go through the entire path of independent training.

This discussion about distillation has thus gone beyond the technical method itself. As China's large models are gradually approaching the global cutting-edge, a more important question is placed in front of the industry: when the catch-up enters deeper waters, what exactly do large model companies compete for?

To understand this divergence, we first need to go back to the specific position of model distillation in the large model industry.

Model distillation itself is a mature technology. Simply put, it uses a more capable "teacher model" to guide a smaller-scale "student model", which can greatly reduce parameters and reasoning costs while improving model deployment efficiency. Today, this is already a common method for the industry to optimize model performance.

The focus of industry controversy actually lies in another type of behavior: using the closed-source model interface of competitors, obtaining output data through large-scale calls, and then using this data to train one's own foundational model. This practice is commonly known in the industry as model "black-box distillation".

The appeal of this model is obvious. For teams with limited resources, directly calling leading models to generate high-quality samples can save massive data construction costs, allowing new models to achieve "overtaking on curves" in hard indicators such as coding, mathematics and logical reasoning. Previously, many open-source communities also used this kind of output migration from "teacher models" to quickly improve the performance of small models.

However, distillation inherently has the "ceiling of the teacher model". The competition of foundational models is ultimately not only a competition of scores on rankings, but also a comprehensive precipitation of data systems, architecture evolution, pre-training, reinforcement learning and engineering capabilities. If a large amount of training data comes from rival models, the company can only learn the capability boundaries of the other party in the end, and it is difficult to touch the real core driving force.

For large model companies that are determined to break into the global first echelon, long-term dependence on the output of external models is tantamount to handing over their own upper limit to competitors. The deeper problem lies in the long-term evolution of R&D organizations.

The competition of large models is essentially a construction project of technical infrastructure. Many core capabilities have to be developed through long-term independent exploration, such as high-quality data governance, training system optimization, algorithm innovation, and reinforcement learning design. If the team is used to relying on external data to improve indicators, the organization will easily degenerate into an engineering method that only focuses on data filtering and fine-tuning optimization, and then lose the ability to explore the underlying laws of models.

At the same time, external model distillation is also facing increasingly severe commercial and compliance risks. Major American large model giants have successively upgraded abnormal monitoring of product interfaces, and restricted data collection and model training in their service terms. Using the output of competing models to train their own products has not only triggered discussions on the R&D boundary in the industry, but also brought many variables at the compliance level of intellectual property rights and service terms.

An incident that caused a stir occurred in February this year, when Anthropic accused DeepSeek, Moonshot and MiniMax of using Claude on a large scale for so-called "distillation attacks", claiming that they generated more than 16 million interactions through tens of thousands of accounts. In July, after Moonshot AI released its new-generation model Kimi K3, it was quickly noticed and accused by American companies. That is to say, model distillation is changing from a purely technical issue to a core topic in intellectual property, service terms and international AI competition.

After all, ByteDance is a large company, and every move of TikTok in the United States is being watched. It went through harsh investigations in previous years, so the compliance risks it can bear are not at the same level as ordinary startups. Zhang Yiming's caution about distillation at this time is, to put it bluntly, that he would rather go slower and pay higher costs than leave the slightest handle to competitors in terms of underlying copyright and compliance.

As American giants continue to tighten API permissions and audit risk control, the technical path maintained by "borrowing strength" is facing more uncertainties. Reducing dependence on external data supply is not only to deal with the current compliance and legal risks, but also ByteDance's advance planning of its R&D system for future participation in global market competition.

Organizational Mechanism for Preserving Creative Capabilities

Zhang Yiming's new judgment on the large model path is changing the R&D method of the Seed team.

Giving up dependence on external model distillation means that Seed needs to bear higher uncertainty. Without mature references, the R&D team needs to explore on their own how to construct training data, how to adjust the model architecture, and where capability breakthroughs come from. The R&D of foundational models is naturally accompanied by a large number of failed experiments, which truly tests the organization's tolerance for long-term exploration.

A large-scale pre-training usually takes several months, and new directions may invest huge resources but cannot immediately produce obvious results. If the evaluation system relies too much on short-term rankings and phased results, it is difficult for the research team to continuously explore unknown paths.

Recently, multiple media reports stated that ByteDance's Seed is advancing the R&D of an ultra-large-scale foundational model. The Financial Times reported that ByteDance is training a model that may reach the scale of trillions of parameters, which is still in the R&D stage, and the final parameter scale and release time have not yet been determined.

The ultra-large-scale training plan sends a clear signal that ByteDance hopes to continue to expand the pre-training scale to find room for model capability improvement. Seed's exploration in the video model field has greatly enhanced ByteDance's confidence in the underlying training path.

In February this year, ByteDance successively launched the video generation model Seedance 2.0 and the image generation model Seedream 5.0 Lite. The two visual models quickly attracted the attention of the industry. Seedance 2.0 was particularly outstanding, which became a hit on social platforms at home and abroad during the internal test phase, and became one of the most discussed video generation models at that time.

The popularity was quickly converted into revenue. According to 36Kr's report in June, Seedance 2.0 has become the most important growth source of Volcengine's MaaS business, contributing more than 1 billion yuan in monthly revenue. Volcengine also raised its 2026 MaaS business revenue target to 15 billion yuan in April, while the full-year MaaS revenue in 2025 was about 1.5 billion yuan.

After the competition of large models enters the deep zone, the organizational method also faces adjustments.

For a long time, ByteDance has relied on a multi-domain horse-race mechanism to promote innovation. Multiple small teams find opportunities through rapid trial and error, and then invest resources according to the results. This method is suitable for product competition at the level of Internet applications, but it cannot fully adapt to foundational model R&D. One training session requires a large amount of computing power, and parallel exploration in multiple directions will bring significant resource pressure.

Therefore, Seed is increasing the resource concentration of core model R&D. LatePost reported that ByteDance is integrating coding-related R&D resources and attracting technical talents including researchers from DeepSeek to join Seed to strengthen the layout in related directions.

Organizational capability has become an important variable in large model competition. Foundational model R&D involves multiple links such as computing power, data, algorithms, and talent systems, and no single point breakthrough can determine the final result. The R&D team needs to face long-cycle investment and bear the uncertainty brought by exploration failures.

Gartner predicts that by 2028, about one-third of enterprise software will have AI Agent capabilities, and more than 60% of generative AI applications will be deployed through enterprise-level AI platforms. For commercial businesses such as Volcengine and Doubao, the underlying model capabilities determine the future competition space. Application-layer products can iterate quickly, while foundational model capabilities require long-term accumulation.

In fact, Zhang Yiming's most fundamental action logic is to maintain the organizational mechanism that generates creative capabilities. From this perspective, a series of recent actions by ByteDance around models, organizations and R&D systems all point to the same goal — to keep the key variables that determine technological breakthroughs inside the company as much as possible.

This line of thinking has a very strong platform nativism feature. The most important capability in the platform era is control: control over algorithms, traffic, data, distribution, and key nodes in the value chain. In the AI era, this methodology continues to extend to model R&D. Models, data, computing power, talents and products need to be integrated into the same system. Long-term dependence on external parties in any key link will weaken the overall independent evolution capability.

Therefore, what Zhang Yiming is really alert to may be that R&D capabilities are led by external technical paths. Distillation can narrow the gap, and APIs can quickly make up for capabilities, but long-term dependence on external momentum will gradually make ByteDance's internal R&D system lose the ability to find the next breakthrough.

For ByteDance, the real moat comes from the R&D machine formed by the combination of talent density, resource volume, organizational efficiency and long-term investment. This is also one of the reasons why ByteDance is known as the "App Factory" in the industry.

Looking back at ByteDance's past battles, the whole set of tactics is familiar and proficient to Zhang Yiming. Systematize key elements, organize and mobilize resources, and then turn phased backwardness into its own advantage of catching up later through continuous iteration.

Facing the key judgment of this meeting in ByteDance's AI competition, what Zhang Yiming wants to maintain is the system that "builds the cannons". Once the next technological breakthrough occurs, ByteDance can still make the cannons by itself.

Diverse Paths in Large Model Competition

The discussion around model distillation, when explored in depth, has touched on the path differentiation after China's large model industry enters the next stage.

In the past few years, the common task of Chinese model companies was clear: to narrow the capability gap with the world's cutting-edge models as soon as possible. Open-source models, synthetic data, model distillation, and engineering optimization have all become important tools, and Chinese teams have thus shown very high catch-up efficiency.

Stanford University's 2025 AI Index records this process. At the end of 2023, the gaps between top Chinese and American models in MMLU, MMMU, MATH and HumanEval reached 17.5, 13.5, 24.3 and 31.6 percentage points respectively. By the end of 2024, the four gaps had narrowed to 0.3, 8.1, 1.6 and 3.7 percentage points. By March 2026, the comprehensive performance gap between top Chinese and American models had further narrowed to 2.7%. Both Alibaba and DeepSeek have entered the sequence of the world's top models.

Behind the catch-up speed, there is a more noteworthy resource contrast. In 2025, private AI investment in the United States reached 2859 billion US dollars, while that in China was about 124 billion US dollars, with the former being 23 times the latter. In the same period, American institutions launched 59 representative AI models, while China launched 35. The gap in resource investment is still huge, but the gap in model capabilities continues to narrow. Engineering efficiency has thus become one of the most distinctive competitive capabilities of China's large model industry.

What I care more about is how Chinese model companies will answer the next question after the capability gap narrows.

Where do the next-generation capabilities come from?

The answers have shown obvious differentiation.

DeepSeek has taken an efficiency path with rich local characteristics. DeepSeek-V3 adopts methods such as MoE, MLA, and FP8 training to optimize the model architecture, training system and hardware utilization in the same engineering system. The 671-billion-parameter V3 only activates 37 billion parameters per Token, and the complete training consumes about 2.788 million H800 GPU hours. DeepSeek then distilled the reasoning capability generated by R1 into small models, and explicitly allowed developers to use model outputs for further fine-tuning and distillation. Technological breakthroughs, cost reduction and open diffusion can occur simultaneously in the DeepSeek system.

Alibaba's Qwen is another path. Qwen continuously promotes the capabilities of foundational models, while continuously expanding model sizes, open-source versions and tool systems, and then connects the models to Alibaba Cloud, Model Studio and the enterprise Agent ecosystem. By 2026, Alibaba Cloud has integrated models, AI infrastructure, development platforms and Agent products into the same enterprise service system. The role of Qwen has gradually gone beyond a single model, and it is more close to a public base layer of Alibaba's AI ecosystem.

Kimi chooses to invest limited resources in several clear directions. Long context, coding, complex tasks and Agent have long maintained high priority. Kimi K3 released in July 2026 has 2.8 trillion parameters, supports 1 million Token context, and focuses on long-cycle programming, knowledge work and complex reasoning. The resource constraints of startups are more obvious, so the technical path needs to be sharper, concentrating resources to find capability fulcrums that can establish leading positions.

Observing Zhang Yiming's choice in such an industrial coordinate system, the logic becomes gradually clear.

ByteDance is taking a path with strong platform nativism, the core of which is to control the key variables that generate capabilities. Looking at Zhang Yiming's more than ten years of entrepreneurial experience in the past, this logic has strong continuity. In the platform era, ByteDance put algorithms, data, traffic and products into a closed-loop ecosystem, forming competitive advantages by controlling key variables. Entering the AI era, the scope of control continues to extend to models, computing power, talents and basic research.

In the next stage of large model competition, the core issue will return to the source of capabilities. Whoever can continuously discover new growth paths for model capabilities will have more opportunities to obtain long-term initiative.

Therefore, in Zhang Yiming's judgment, distillation involves not only model performance, but also where the core production capabilities come from. Once the core production capabilities come from competitors for a long