2 billion RMB, Alibaba leads the investment in a post-90s intern: working as a teacher for AI
An AI company established less than a year ago is securing a 2 billion yuan financing round.
According to reports, Alibaba plans to lead a new round of financing of approximately 300 million U.S. dollars for UniPat, with Tencent, HSG and other parties participating. The company is valued at about 2.5 billion U.S. dollars, equivalent to nearly 17 billion yuan.
UniPat designs questions and grades papers for AI, telling the model where it made mistakes.
This business has already started making huge profits.
In 2024, AI data and evaluation company Surge AI recorded a revenue of about 1.2 billion U.S. dollars; Mercor disclosed earlier this year that its annualized total revenue has exceeded 2 billion U.S. dollars; Scale AI's revenue last year also approached 1 billion U.S. dollars.
- 01 - Three billion-dollar businesses emerged in just one year
UniPat was founded at the end of 2025. Its founder Li Kuan, a post-90s generation, graduated from Beihang University with a bachelor's degree, and obtained his master's degree from the Institute of Computing Technology of the Chinese Academy of Sciences. He is currently pursuing a doctorate in the Department of Computer Science and Engineering of the Hong Kong University of Science and Technology, with research directions including LLM Agent. Before starting his business, he participated in post-training research as an intern at Alibaba's Tongyi DeepResearch, focusing on data synthesis and reinforcement learning.
One of the core problems solved by post-training is to make the pre-trained model continue to "become smarter". It is not enough for the model to be able to solve a certain number of problems. The training team also needs to know: what counts as a good answer, where the mistake lies, and what kind of feedback is most worthy of being used for further training.
UniPat has turned this matter into a product. A person in charge of the AI agent project at a leading large technology company told Pencil News that UniPat replaces the work originally done by the test team in the agent R&D project, packaging the entire evaluation process into a product to assess the perfection of AI task execution, which hits a rigid demand.
At present, the most direct way to make money is to sell the judgment of professionals to large model companies.
Mercor organizes programmers, lawyers, financial practitioners and scientists for clients including OpenAI, Anthropic and Google DeepMind, asking them to design questions, write answers and evaluate model outputs, and then deliver the data to clients for post-training. The company initially focused on AI recruitment, but later found that the parties most willing to pay for high-end talents are large model laboratories, so it shifted its business to AI training.
According to Forbes, Mercor's annualized total revenue in June this year exceeded 2 billion U.S. dollars, doubling from 1 billion U.S. dollars in February. The company's total revenue in the first half of this year was about 614 million U.S. dollars, of which about 91% came from foundation model companies.
However, the 2 billion U.S. dollars is not pure software revenue. About 60% to 70% of the cash flow is paid to experts. Mercor earns the price difference from matching, and more importantly, it can quickly find suitable experts and control delivery quality.
Handshake follows the same path. It was originally a college student recruitment platform with a network of universities, students and alumni. After entering the AI training track in 2025, the annualized total revenue of this business has approached 1 billion U.S. dollars in less than a year; after deducting expert remuneration, the annualized net revenue is about 300 million U.S. dollars.
Surge AI is more like a high-margin data service provider. Its founder Edwin Chen once worked at Google, Facebook and Twitter. The company's revenue in 2024 was about 1.2 billion U.S. dollars, and its clients include Google, Meta, Microsoft and Anthropic.
Surge also sells human feedback and high-quality data, but it pushes its business to higher value links: instead of just finding people to solve problems, it designs tasks, scoring standards, evaluation sets and training environments. Compared with pure manpower matching, this is closer to a complete set of AI post-training services.
These companies are earning the same type of revenue: the stronger the model is, the more it needs more professionals to tell it where it went wrong.
- 02 - Good answers are extremely valuable
The reason why the industry is shifting from labeling to evaluation is very simple: models are getting better and better at taking exams.
The problem of early training data was "whether there is an answer". Today's problem has become "what counts as a truly good answer". To make AI recognize a cat, one label is enough; but to make AI modify a large code repository, produce an investment banking analysis, or complete hospital administrative procedures, it is difficult to score only with the two words "right/wrong".
At this time, the scoring standard itself becomes data.
ExpertEval launched by UniPat this year covers three fields of healthcare, finance and law, with a total of 3213 expert cases and 207 scenarios, with an average of 21.83 scoring criteria set for each case. It not only checks the conclusion, but also verifies the evidence, reasoning process and high-risk errors. These scoring results can be continuously used for post-training.
Its Monthly-SWEBench draws questions from the latest closed GitHub tasks every month to test the model. The question bank must be updated continuously, because public benchmarks will soon be "crammed": higher scores do not mean that the real working ability is improved synchronously.
Its SaaS-Bench is closer to real work scenarios. UniPat puts the Agent into 23 sets of operable SaaS systems, requiring it to complete tasks such as reimbursement, account closing, project management, and medical administration. The results show that although the best-performing model can complete many intermediate steps, the proportion of completing the entire task end-to-end is only 3.8%.
This 3.8% is exactly the business opportunity for such companies. If the Agent can stably complete 99% of the work, the value of the evaluation company will instead decline. It is precisely because the model fails frequently that AI laboratories need to constantly find new questions, identify errors, design scoring rules, and then send the failed samples back to training. Every time the model is upgraded, this cycle may repeat.
Scale AI has also gone through a similar upgrade. It started with autonomous driving data labeling, and later expanded its business to high-quality training data, model evaluation, red team testing and government projects. In 2025, Meta spent 14.3 billion U.S. dollars to acquire a 49% stake in the company, raising its valuation to 29 billion U.S. dollars.
The value of such companies lies in their ability to continuously discover things that the model cannot do, and then turn the cause of failure into training signals. Every time the model is upgraded, new evaluation and training demands will emerge again.
- 03 - The next round of revenue
The industry is looking for a better business model than "selling data by head count". The new direction is called RL Environment, that is, reinforcement learning environment.
It can be understood as building a "simulated company" for the Agent. There are emails, CRM, code repositories, financial systems, customer work orders and various tools in it. The AI does not just answer one question, but continuously completes dozens or even hundreds of steps of operations. The system then automatically scores according to the results, allowing the model to practice repeatedly.
This is more difficult than traditional labeling, and closer to software products. The traditional model requires continuous recruitment, order assignment and acceptance, and revenue growth is often accompanied by the growth of labor costs; once the training environment is built, it can be called repeatedly by different models, and continuously generate new training trajectories and feedback signals.
Surge has established a dedicated team to build such environments. This year it released EnterpriseBench CoreCraft, which simulates a customer support company with more than 2,500 entities and 23 tools, allowing the Agent to handle multi-step work. Mercor also acquired Deeptune, a company dedicated to reinforcement learning environments, in July, and clearly regards training environments in fields such as coding, healthcare, and law as its next stage of business.
Handshake has started to sell another scarce resource: real enterprise operation data. It publicly solicits desensitized business data from enterprises, paying 100,000 to more than 4 million U.S. dollars for each cooperation according to the scale and value of the data, and then provides the sorted data to cutting-edge AI laboratories for training and evaluation.
This shows that this business is forming three layers of charging: the first layer sells expert time, the second layer sells sorted training data and scoring standards, and the third layer sells environments that can continuously generate training data. The further you go, the closer you get to infrastructure.
Of course, this business has not got rid of human services yet. About 60% to 70% of Mercor's cash flow needs to be paid to experts, and its clients are highly concentrated in a small number of leading AI laboratories; after Meta took a stake in Scale, it experienced changes in some customer relationships, which also shows that the "neutral third-party" attribute itself is a kind of asset.
This article does not constitute any investment advice.
This article is from the WeChat official account "Pencil News" (ID: pencilnews), written by Song Ge, edited by Huang Xiaogui, and published with authorization from 36Kr.