HomeArticle

An intern who once worked for Kimi founded a company, which is now valued at nearly 20 billion yuan.

36氪的朋友们2026-09-14 15:45
The first bucket of gold for young people.

For a long time, large model data companies have not been in the field of vision of VCs, but this situation is changing. I recently heard that several young people with internship experience at leading model manufacturers such as Alibaba's Tongyi and Moonshot AI have founded a data company, which is highly favored by capital.

This company takes AI training data as its entry point, and has won tens of millions of dollars in orders in less than a year, completed about three rounds of financing, and its valuation has also risen steadily. The latest foreign media news shows that this company is about to complete a $300 million financing round, with a star-studded investor lineup including HSG, Alibaba and Tencent, and its valuation reaches $2.5 billion.

Of course, there were not no data service companies in China before, but before this year, such companies were rarely noticed by venture capital. The reason is that this model is mainly a cash flow business, and the entrepreneurial threshold in the AI field is not high. Earlier, a number of data companies in third- and fourth-tier small cities used cheap labor to provide corpora for model training. With the improvement of the intelligence level of large models, the requirements for data quality have increased simultaneously, and more and more highly educated practitioners from top universities have joined. Senior professionals from all walks of life, such as legal workers, doctors, journalists, screenwriters, etc., have begun to settle on the platforms of some data service companies. As a part-time job, the income of data labeling ranges from 500 to 1000 yuan per hour.

However, if it only operates as an expert network, it is difficult to reach a valuation of tens of billions of yuan in China. The data-selling model is still a business of phased sustainability in the eyes of many investors.

An investor who has been in contact with the aforementioned project told me that before the $200 million financing round, investors placed bets based on their judgment of the founders. Now, to raise financing at a valuation of $2.5 billion, this company must tell a new story. As far as I know, in the first half of this year, it has already set foot in the financial sector, built a comparable and verifiable system in the prediction market, and it is said that it will make profits by selling APIs.

There are about more than 20 domestic data training companies, and the aforementioned cross-border financial story is only one of the future directions they try to describe. After communicating with investors, I found that traditional data service models that are difficult to tell sufficiently attractive growth stories, such as manual labeling, expert outsourcing, and data delivery, have new interpretations among more and more young post-95s and post-00s entrepreneurs. Relying on their experience in model training and sufficient understanding of model capabilities, they have made a originally relatively boring cash flow business begin to have new room for imagination, thus stirring up a wave of financing enthusiasm.

New Demands for Data Are Emerging

To understand the business run by these interns, we may start from the hot topic in the media last year that "postgraduates and doctors from 985 and 211 universities go to be AI trainers". For example, graduates from the Chinese Department of Peking University working at DeepSeek, and medical doctors working at Kimi. Some of them mocked themselves online as "data foremen", leading a number of labelers in the form of BPO (outsourcing) to cooperate with the R&D team to produce training data for models.

The business logic of selling data is not complicated. As long as orders can be obtained, the company can run into a company with good cash flow, and it does not necessarily need financing. Surge AI is a typical case. Founded in 2020, the company's founder Edwin Chen previously worked on machine learning at Google and Meta before starting his own business. This company did not take external financing for a long time and mainly developed on its own. Until 2025, the market began to spread news that it was seeking up to $1 billion in financing. At that time, the company's revenue in the past year had exceeded $1 billion, and the valuation discussed in the market also exceeded $15 billion.

With the improvement of model capabilities, the demand for labeling by senior talents in professional fields such as medical care, law, and finance, as well as complex tasks that require subjective judgment, is growing rapidly. High-quality human labelers are scarce in the market, especially professionals in professional fields, and large AI companies are willing to pay a high premium for this. Mercor seized this opportunity. This company was originally an AI recruitment platform that provided "contract talents" for large data labeling companies in the early stage, and later directly organized professionals into a training data supply network. In 2025, Mercor raised financing at a valuation of $10 billion, and in 2026, it was reported that it planned to raise about $5 billion at a valuation of $20 billion.

In short, many unicorns have been born overseas, but few domestic AI training data service providers have the opportunity to obtain VC support.

An investor who has investigated these model factories told me that at that time, the internal standards for data procurement, including the standards for new suppliers, were not particularly clear. Although they have cooperated with several Internet manufacturers and even some overseas suppliers, the overall supply volume is not that large. On the other hand, for a period of time in China, distillation was the main focus, which largely weakened the necessity for everyone to find expert data.

This year, Agents have exploded in an all-round way, and the type of data required is no longer simple expert data, but data that is closer to the real environment.

A researcher at a model factory told me that the data needed now is no longer just "questions and answers", but a closed-loop, verifiable training environment that allows the model to act repeatedly and obtain feedback. The model performs tasks in it, the environment records every step of its actions, then the validator judges the result, and finally returns the feedback to the model. For example, if you want to train a model to write better front-end code, it is not enough to manually tell it "whether this code is good or not". You need an environment that can execute code, and a model that can understand the page effect to act as a referee, and then run for hundreds of rounds or even longer to continuously give feedback to the model.

Some American startups have begun to specialize in reinforcement learning environments. AfterQuery hopes to let models and agents learn to complete tasks like professionals. The company describes it as "encoding the patterns, decisions and reasoning of the world's top practitioners", and outputs expert reasoning datasets, reinforcement learning simulation environments and model evaluation services for cutting-edge large models.

On the demand side, the management of Anthropic has also discussed investing $1 billion in reinforcement learning environments within the next year. For OpenAI, the total data expenditure in 2025 was about $1 billion, and the internal forecast will rise to $8 billion by 2030. Now there are almost a dozen such seed-stage teams in the United States, with no more than 20 people, serving 1 to 3 large customers.

Many such entrepreneurial opportunities have also begun to appear in China. An investor told me that more and more researchers are spilling out from model factories. By May and June this year, this trend became relatively clear: a number of new teams providing high-quality data began to appear in the market, and at the same time, the demand of model factories for data in new fields is also rising.

Post-00s Are Flocking to Wealth Creation

Like entrepreneurs in other fields of AI, the founders of training data companies are usually very young. Alexandr Wang, the founder of Scale AI, became a billionaire at the age of 24, and the three founders of Mercor also joined the ranks of billionaires at the age of 22 with a valuation of tens of billions of dollars.

The latest example is the two co-founders of AfterQuery. Spencer Mateega is 23 years old, and Carlos Georgescu is 22 years old. Both of them were still in college when they founded the company. Their company just announced that it is raising a new round of financing with a valuation of $3.2 billion. Founded only 18 months ago, this company's main business is AI training data services.

Nowadays, the founders of domestic data startups are the same, and they are not necessarily the "little geniuses" in the embodied track, just like the interns from leading large model manufacturers I mentioned at the beginning.

An investor who has been in contact with a domestic unicorn data training company told me that in the early stage, everyone's judgment was still more focused on the "people". People with both business sense and researcher attributes are still relatively rare in the market today. The founder of this company had internship experience in first-tier large model manufacturers such as Moonshot AI, participated in model training work and published papers.

The current situation is that the front-line work of model training is mainly undertaken by PhDs from top universities, and they will undertake specific execution work since their internships. After they complete these work with the outsourcing team, it is easy for them to think that since this matter can be completed in the outsourcing department of a large factory, why not do it on their own? As a result, some people began to come out of model factories or outsourcing teams to set up data startups.

This is also the common feature of this batch of data entrepreneurs: they are closer to the actual needs of model factories and more sensitive to changes in the external market. An investor who has invested in similar projects told me that his judgment on the portrait of entrepreneurs is mostly that the team must at least reach the T0 or T1 level, and it is best to be the post-95s, post-98s to post-00s generation who have actually participated in front-line model training.

Engineering capabilities are also critical. How to build a team and how to implement the data quality requirements recognized by researchers into actual production requires strong engineering capabilities. From the construction of the entire data pipeline, to Query design, initial data composition, to subsequent cleaning and quality control, there are complex engineering links behind it, which is not simply manual labeling.

Data seems to be a business on the surface, but when done well enough, it may also become a candidate team outside the model factory, or even a small pre-training and post-training organization. Investors can obtain a team with high researcher density at a small cost in the early stage. The data business is a way for them to support themselves in stages, and it can also help them prove their position in the industry and keep up with the changes in the development of model intelligence.

The Ceiling of the Data Business

However, a pure data business can hardly support these startups to move forward sustainably.

In 2025, the market disclosed that Mercor, which started with expert data in vertical fields, achieved an annualized revenue of $100 million, but 60% to 70% of the total revenue needed to be paid to contractors, leaving very limited space for the company. More than half of Scale AI's revenue is also used for direct business costs, including contractor salaries. Among the prices announced by some domestic data labeling platforms, the hourly wage of labelers is usually between 200 and 500 yuan, and those with higher qualifications and stronger professionalism can even reach thousands of yuan.

As model capabilities continue to improve, the demand of post-training companies for top experts will become more and more complex. If they still rely on a large number of manpower to find experts and produce data, it means that costs are rising continuously; if the fees charged by experts are increased, gross profit will be compressed; if the fees are not increased, it will be difficult to obtain sufficiently high-quality data.

Moreover, data business orders are not completely certain. A friend of mine who is an investor told me that when the first batch of data suppliers entered the market, many companies first produced a batch of high-quality samples in each vertical field, whether it was manually polished by hand or produced through other methods, they would put a lot of effort into the samples, and then show them to customers such as Alibaba and Tencent. After the customer approves, they will place further orders. But in the data industry, one order does not equal a sum of revenue that can be confirmed immediately. Each batch of data needs to be inspected and quality controlled by researchers. If the quantity or quality of the delivered data fails to meet the standards for several consecutive times, the customer may completely reduce or even cancel the order. For example, a customer signing a $100 million order does not mean that the company will definitely be able to achieve the corresponding scale of revenue in half a year, one year or two years later.

From the perspective of VCs, investment ultimately returns to two questions: can this company form a sufficiently high value multiple, and is there a clear exit path in the future?

The division of labor among overseas large factories is relatively mature. The annual budget allocated by each large factory to external data companies will still grow by 2027 and 2028, and the volume is very large. Therefore, overseas data companies do have good market space, and there have been exit opportunities through mergers and acquisitions. But the logic of the Chinese market is different. If it is just a data outsourcing company in the end, whether it is merged and acquired or listed, the space is relatively limited.

In the view of the aforementioned investor, the future data business may no longer be the current model of "how much money to sell for one piece of data", and some data startups are beginning to try to stand at the SaaS level. For example, some traditional enterprises do not have the ability to train AI models, but they have a large number of unique business processes and high-value data internally. Data companies can enter the workflow of these enterprises to help them complete AI transformation.

In this process, enterprises will open some unique business scenarios. The company not only obtains data, but also accumulates real task trajectories and interactive feedback, and then precipitates these things into its own model and training system. In this way, data is no longer just a one-time delivered commodity, but has become a layer of infrastructure that connects customer workflows, model training and long-term services. The pricing method of data companies will also change. Large manufacturers will no longer simply pay according to the number of data pieces, but pay for a system that can continuously generate high-quality data, complete task execution and model feedback.

Some overseas data companies have already moved towards this deeper service capability. Companies such as Turing, Scale, and Invisible are developing consulting or deeper enterprise service businesses, hoping to further intervene in customers' AI workflows from simply providing data and manpower.

Of course, the deeper the data goes into the core business of customers, the higher the risk. After Mercor encountered a data leakage incident this year, Meta suspended its cooperation with it, and OpenAI also launched an investigation into it. The relevant data may involve the training methods of model companies, contractor information and proprietary data.

Where Is the New Growth Story

In addition to the pursuit of higher intelligence, the accelerated release speed of model versions is another reason for the growth of data orders this year. Since 2026, the release rhythm of leading large model manufacturers has been compressed to a monthly level. From a technical perspective, there may be a key variable playing a role behind this: RSI.

RSI (Recursive Self-Improvement) refers to the fact that AI participates in and even improves its own R&D process, thereby continuously shortening the model iteration cycle. It is regarded as an important driving force for the acceleration of model releases in this round, and has also become a hot word in the AI venture capital circles in China and the United States.

This is not a brand new concept. The academic circle has discussed it very early, and the industry has also been trying to promote it. In May, Recursive Superintelligence (RSI), a new AI laboratory co-founded by 8 top AI researchers including Tian Yuandong, former research director of Meta FAIR, completed $650 million in financing, with a post-investment valuation of about $4.65 billion. This technical concept was officially pushed to the capital market.

Some data startups have also begun to take RSI as part of their financing narrative.

My investor friend noticed that some data teams he has looked at recently have begun to build "semi-RSI" systems in new fields. Of course, they have not achieved fully automated iteration, but in some links, they may have reached close to half, or even 70% to 80%. For example, the team can quickly produce high-quality analysis, labeling and synthetic data in a new vertical field, and at the same time efficiently embed expert experience and human judgment into the system. In the future, if manual intervention can be further reduced, allowing the system to complete task generation, execution, evaluation and iteration on its own, there is an opportunity to move from data production to a more complete model improvement closed loop.

However, RSI entrepreneurship is a game with extremely high thresholds. It requires cutting-edge models, huge computing power, top research talents, training infrastructure, large-scale experimental capabilities, and a sufficiently strong model evaluation system at the same time. At present, the parties that truly have the ability to try to close the complete loop are still mainly OpenAI, Anthropic, Google DeepMind, and very few new Labs with extraordinary capital and talent allocation.

Many companies in the market today are talking about RSI, but what they are doing is not the same thing. Some teams are benchmarking against large model companies, aiming to develop the next generation of models and try to run through the complete model training process; some RSI only occurs at the automation level of a certain product; what some other teams call RSI is that their own data pipeline can achieve self-purification and self-iteration. The levels, directions and scopes of these works are different, and there may be very obvious gaps between them in the short and medium term.

Therefore, from an investment perspective, we must first dismantle the specific meaning of RSI: which level of problem is it solving? What kind of ecological niche can this level occupy in the entire industry in the future? Where is its capability boundary, and is there any possibility to continue to extend upward?

Of course,