HomeArticle

From interns earning a daily salary of 5,000 yuan to annual pay packages exceeding 100 million yuan, the competition for AI talents has spiraled into utter madness | DeepKr

王毓婵2026-09-08 16:08
Managing large language models is just like preparing a simple delicate dish, with young prodigies grasping the ladles of industry giants.

Written by Wang Yuchan, Wen Lihong

Edited by Zhang Yuxin, Yang Xuan

Source: Intelligent Emergence (ID: AIEmergence)

Cover Source: AI-generated

Hot Money

When Zhu Kexin learned the annual salary of his 1998-born colleague, a frontline algorithm researcher who had just switched to a top-tier tech giant, he felt dizzy.

The pay rise was as high as 300% — when Zhu Kexin and his colleagues guessed the new pay of their former teammate over a dinner gathering, the most aggressive guess was a 50% increase. But the actual figure turned out to be a jump from 1 million yuan to 3 million yuan, far beyond all their expectations. Within the span of that meal, everyone at the table re-evaluated their own market value.

The old salary benchmarks were rewritten rapidly, and this "dizziness" kept happening. Only one year has passed since that job switch, Zhu Kexin says, and now no one is surprised to hear that a frontline researcher can get an annual salary of 3 million yuan.

The age group that can access such high annual salaries is also getting younger.

A post-2000 graduate student, before even graduating from campus, was paid a daily salary of 5,000 yuan as an intern, with a monthly income far higher than most white-collar workers in this big city. His supervisor told him privately, "You can get an annual salary of 3 million yuan after you become a full-time employee upon graduation." This young man told 36Kr directly that he just wants to graduate immediately this year.

From Silicon Valley to China, when top AI talents are being fought over with annual salaries reaching hundreds of millions of yuan and top executives personally extending recruitment offers — Wu Yonghui, the head of ByteDance's AI division, was recruited by Zhang Yiming after 8 months of negotiations; Meta not only offers sky-high salaries, but Mark Zuckerberg even personally "delivers homemade soup" to OpenAI researchers he wants to poach — across the entire AI talent pyramid, from the top to the base, the compensation for talents at every layer is being pushed higher and higher.

36Kr Illustration

People at the top of the pyramid can get annual compensation packages worth hundreds of millions of yuan, but this group is extremely small, with only a few people on each company's model team.

The next layer down are those who get annual packages of over 10 million yuan at large tech giants. A headhunter told 36Kr that every large-scale company has roughly 10 to 20 such people, each of whom is an industry star whose job changes will draw widespread attention across the sector.

The layer below that includes top fresh graduates such as TopSeed talents at ByteDance, AliStar at Alibaba, and Tencent's "star recruits" with annual packages of 5 million yuan and above, with dozens of such people at each company. A PhD candidate who joined ByteDance this year explained the composition of his 6 million yuan annual salary to 36Kr: "It's not all cash, it also includes ByteDance's stock options, Doubao shares, and bonuses." He also mentioned that some fresh graduates can get an even higher figure than 6 million yuan, "but that's a very rare case."

Further down are frontline AI algorithm researchers with annual packages ranging from 2 million to 3 million yuan, known as "regular workers", the largest group, whose compensation is determined by their published academic papers and core project experience. "This year, an annual salary of over 2 million yuan is the baseline that an ordinary algorithm researcher can accept, while last year the baseline was just over 1 million yuan," an algorithm researcher at a large tech giant said. "I have a friend who got an annual package of over 2 million yuan from Seed before, but he turned it down, because another AI company offered him over 3 million yuan and a higher position."

The very bottom of the pyramid is made up of AI algorithm interns. PhD interns working on core directions at ByteDance, Tencent and Alibaba can get daily salaries of 5,000 to 6,000 yuan, while two years ago, the average monthly salary for PhD interns was 10,000 yuan, which is only equivalent to two days of their current pay.

When the monthly salary of interns has exceeded 100,000 yuan, the competition for AI talents, especially for large model researchers, has entered a white-hot stage across every layer of the talent pyramid.

Behind this is the evolution of competition between Chinese model developers: ByteDance used to be the most generous buyer of AI talents in the market, but in the past year, Tencent has also rushed to recruit talents and become a big spender for top talents. At the same time, DeepSeek and Moonshot AI have completed large financings in the primary market, Zhipu AI and MiniMax have gone public, and leading large model startups have also raised more capital to increase their investment in talent recruitment.

All along, "talent" has been a key variable in competition across all industries. But the intensity of the talent battle in the current AI industry is significantly higher than that of the internet industry, the previous sector that saw massive hot money inflows. The only comparable scenario is probably the scramble for core talents during the early rise of the semiconductor industry between the 1950s and 1970s.

The common feature of these two industries is: companies seem to own the technology, but the people who truly master the technology are those dozens of core technical personnel. Therefore, talent has become the key "bottleneck-breaking" variable in corporate competition.

In this way, competition between AI companies has evolved into competition for talents. And the flow of talents is drastically changing every company involved, and profoundly affecting the pattern of the entire AI industry.

Buying the Recipe

What large tech giants and large model companies are competing for is essentially the technical "recipe" held by core talents.

When training models, researchers use the term "recipe" to describe all aspects of model training: what size the model should be, how to proportion the data, how to select the training path, and what the key tricks are...

A successful model is developed through a large number of trial and error experiences. The people who have gone through all this master this knowledge. This was also the case in the early semiconductor industry: process parameters, yield rates, material properties, and manufacturing tricks can hardly be fully written into patents. Therefore, poaching the right person can help a company obtain this "tacit knowledge", quickly close the generation gap between models, and narrow the gap with competitors.

There is a good reason for the high price of these recipes. The reset cost of model training is extremely high: training a base model or fine-tuning a trillion-parameter model, a complete experiment can cost tens of millions or even hundreds of millions of yuan in computing power. If the training direction or method is wrong, the loss will far exceed the annual salary of several researchers.

More critically, there are very few people who can master the core recipes.

Unlike most of the so-called entrepreneurial experience in the internet era, the "ten-thousand-GPU training experience" valued in the model circle can only be supported by sufficient computing power and data in a small number of top tech giants and top laboratories; and among these few companies, the number of people qualified as "training commanders" is also very small.

"There are only 200 people in the whole of China who can lead large-scale pre-training projects," said headhunter York. A ByteDance HR also told 36Kr that when they started to focus on poaching core researchers in 2024, they made a list that only contained a few hundred names.

When people who master the core recipes flow between companies, the model training knowledge that was originally held by a small number of companies will be "open-sourced" accordingly, transferred across companies, and finally raise the ceiling of the entire industry. This is similar to the scenario: the recipe of a dish was originally mastered by only a few chefs, and the restaurants that employed these chefs seemed to be very strong for a time; but once a star chef changes jobs, the recipe will spread immediately, no restaurant can maintain a sustained lead, but more and more good restaurants will raise the overall level of the whole industry.

36Kr Illustration

In the second half of 2024, Zhou Chang, who was previously in charge of multi-modal pre-training at Alibaba, joined ByteDance. Although ByteDance paid him an annual salary of tens of millions of yuan, Zhou Chang's arrival quickly helped ByteDance achieve breakthroughs in multi-modal capabilities, which formed the foundation of the now viral Seedance model.

Another industry insider told 36Kr that Cheng Ye'an, a former Zhipu AI researcher and core contributor to GLM-4.5 and GLM-4.5V, joined Moonshot AI this year and is listed as an author on the K3 technical report. His main contribution lies in Agentic Coding post-training, which is exactly the strong suit of Zhipu AI.

For a company, when core talents leave, the loss is not just a workforce, but the accumulated knowledge of the entire team — experimental paths, failure records, debugging experience — that is completely delivered to competitors. This kind of knowledge transfer does not even have to wait until the new employee joins the company. Some people start to quietly share how their internal team operates with rival companies as early as the interview stage. The interview itself has become the most frequent channel for technical details to leak out.

If people who master the recipes choose to start their own businesses, they will definitely become the darlings of capital.

Liu Yu, former research director at SenseTime, went through the full closed loop from data processing, training infrastructure, model R&D to product implementation in 2024 when multi-modal technology was just starting out, and he scheduled more than 4,000 GPUs to train models. At that time, talents with thousand-GPU training experience were already very scarce. So after he left to start his own business, investors flocked to him immediately. Because so many investors came to him, Liu Yu even had to turn most of them away, which made the industry call him "a person that people in the investment circle can hardly meet".

The flow of talents itself is an important part of the capability competition among domestic large model companies in this round. "The core of model competition is mainly computing power, data and people," headhunter Sam told 36Kr. "Only when you find the right people, can the computing power and data be used effectively. Otherwise, no matter how much money you spend, it won't work."

This is why "people" have become such a key asset to compete for in the AI industry.

Of course, in addition to finding that small group of people who master the recipes, large model companies also need to recruit a large number of smart and reliable young people to do the supporting basic work. The talent battle thus broke out in full swing.

Unconventional Tricks

In earlier years, the highland and source of outflow of AI talents were Baidu and Alibaba, which were one step ahead in technology. But recently, ByteDance, which has extremely abundant AI talents, has also become a new "targeted poaching spot".

An insider at a large model manufacturer revealed to 36Kr that they are benchmarking against ByteDance's data system and poach people on a position-for-position basis, "as long as they accept the offer, their salary will be doubled."

ByteDance responded quickly. At the end of 2025, ByteDance sent a company-wide email announcing that its investment in bonuses and salary adjustments would be increased by 35% and 1.5 times respectively, and a new job level system would be launched to raise the salary ceiling for each job level. The adjustment related to job levels is considered a measure specifically targeting AI talents: high-paid researchers cannot be accommodated in the original job level framework.

The measures to retain talents have continued to increase this year: ByteDance has given clear pricing and repurchase terms to "Doubao shares" that are only issued to employees in AI-related businesses, and people who hold Doubao shares can also choose to convert more of their year-end bonuses or total compensation into Doubao shares. This is as valuable as the stock options at the beginning of ByteDance's establishment. These measures may all share a common trigger — nearly 70 key technical backbones left Seed in the whole of last year.

Tencent is not only willing to pay high salaries, but also willing to offer high positions. Star researchers from Silicon Valley such as Yao Shunyu and Tian Yonglong can directly lead a team after joining Tencent, and even become the head of the entire model team. With such examples, headhunters have their own sales pitch when poaching people: "You may just be one of hundreds of top researchers at OpenAI, but if you come to a Chinese tech giant, you can become the top leader of the business line."

However, money and positions are not everything. Many startups have their own "unconventional tricks" for recruitment.

Two AI startups in Shanghai often target the same candidate at the same time. This year, both of them took a fancy to a doctoral student. He had originally prepared to accept an offer from one of them, but was finally poached by the other startup backed by a university. The reason was not money, but the other party said it could provide him with a faculty position at a top Shanghai university.

Zhipu AI is probably the pioneer of this "university track". The company's founder and chief scientist Tang Jie, and almost all the core management of the company come from Tsinghua University. In this strong "academic atmosphere", many Tsinghua students choose to join Zhipu AI as a stepping stone to pursue a master's or doctoral degree under Tang Jie's supervision.

If they cannot provide such strong resources, HR or headhunters will adopt "psychological tactics".

The company that Li Qing works for is one of the "Six Little Tigers of Large Models". Recently, he has received several headhunter calls every week. The calls usually come during working hours, and the other party uses "psychological tactics" — they are not in a hurry to talk about positions and salaries, but focus on negative topics such as the stock price of Li Qing's company, work intensity, and his health status, to shake Li Qing's confidence in his current job from the side.

Li Qing remembers a headhunter who was keen to talk about the stock of his company every time they talked. "Your company's stock price has fallen, do you want to consider other opportunities? XX will definitely have greater potential after going public."

Large tech giants have a more systematic method, which we call "saturated staffing".

An employee at ByteDance told 36Kr that in some core research fields, Seed will definitely maintain two teams with similar capabilities beyond the normal resource allocation. "So even if one team is poached entirely by a competitor, the other team can fully fill the gap. Even if you leave a very core position, you won't cause any waves at ByteDance."

The headhunting industry can also feel this kind of "saturated staffing". A headhunter told 36Kr that among the companies he serves, Tencent will focus on recruiting talents during the team building period, but ByteDance's poaching pace has almost no fluctuations. "No matter whether there is a vacancy or not at the moment, ByteDance will keep recruiting people continuously."

There is also the "radical poaching tactic that cuts off the opponent's supply".

A person from a large tech giant told 36Kr that they are "targetedly incubating model talents from rival companies". "We have identified the 10 most important people of our competitor, and we will try to poach all of them. It would be best if they are willing to switch jobs to us; if they are not willing to come, we will encourage them to start their own businesses, and we can help them find investment."

According to his description, this operation has a clear name list and quota: for example, ByteDance identifies the 10 most important people of Seed, maps their information one by one, and then tries to get in touch with them. Job hopping is Plan A, investing in their startup is Plan B, both plans can achieve the purpose of weakening the competitor.

VCs call this practice "pre-seed negative