The AI assessment company founded by ex-employees has closed a round of funding led by Alibaba at a $2.5 billion valuation.
It is reported that Alibaba will lead a $300 million financing round for UniPat AI, valuing the company at $2.5 billion. Investors including Tencent and HSG may also participate in the deal, but negotiations are still ongoing and the specific terms have not been finalized.
Compared with the high profile of the six rising AI startups, UniPat seems to be relatively unknown to the public.
However, people familiar with the AI industry know that some time ago, after the US venture capital institution Dimension Capital visited China, it met with a number of AI companies intensively and put forward 13 judgments on China-US AI competition. After returning to the United States, they sent an internal letter to their limited partners, which was later widely circulated in the investment circle.
The judgment about data labeling companies in the letter takes Chinese startups represented by UniPat as typical examples.
It is mentioned in the letter that US data companies such as Mercor, AfterQuery and Turing have become suppliers of Chinese large model companies. At the same time, a number of new companies have emerged in China that are no longer satisfied with traditional data labeling, but have started to work on model evaluation, real-world tasks and verification infrastructure.
This enterprise designs real scenarios for testing and measuring artificial intelligence models, and specifically generates detailed training and benchmarking data. The benchmarks it has released so far have been cited by US laboratories in their own model reports.
In essence, this line of thinking still follows the logic: Use higher-quality supervision to replace more computing power. The vast majority of them have been established for less than 24 months, and a few of them have quickly exceeded $100 million in revenue.
The "Chinese Examiner" for AI Model Capabilities
In August this year, Forbes reported that the cooperation between US data companies and Chinese AI companies has formed a considerable scale of data supply chain.
Citing buyer communications and internal platform materials it reviewed, the report said that AfterQuery has a cooperative relationship with Ant Group, Alibaba has previously purchased data from AfterQuery and Mercor, and Turing has also cooperated with ByteDance. The relevant Chinese companies did not respond, and some data companies refused to disclose customer information.
These companies are not traditional "labeling outsourcers".
Mercor focuses on expert data and enterprise workflows, AfterQuery focuses on transforming experts' judgments, trade-offs and working processes into training data, while Turing integrates data, reinforcement learning environments and model evaluation into the same set of services.
This shows that the data industry is undergoing an upgrade.
In the first stage, data companies sell human working time.
In the second stage, data companies sell higher-quality answers.
UniPat is an AI evaluation company that designs real scenarios for testing and measuring artificial intelligence models, and specifically generates detailed training and benchmarking data. The benchmarks it has released so far have been cited by US laboratories in their own model reports.
Ranking on leaderboards is usually an important indicator to measure the technical capability of a large model enterprise. For example, Kimi K3 has demonstrated its advantages through rankings on global benchmarking websites for artificial intelligence analysis and other related fields.
However, with increasingly strict copyright and privacy restrictions, large language models are rapidly depleting the human-generated data available on the Internet, and developers often encounter situations where the scores on leaderboards are artificially inflated and cannot reflect actual performance. UniPat is designed to solve these two bottlenecks at the same time.
At present, US competitor Scale AI (Meta has spent $14 billion to acquire its shares) focuses on data labeling, while Mercor (which is currently in talks for financing at a valuation of $200 billion) operates a data labeling and testing market composed of human experts.
In terms of benchmarking, Artificial Analysis and Chatbot Arena are in leading positions.
Providing Supervision for Chinese Models Facing "Chip and Computing Power Shortages"
UniPat was founded at the end of 2025, and its founders used to work at Alibaba, Moonshot AI and Tencent. During their time at Alibaba, they worked at the Tongyi AI Lab, participating in post-training analysis, data synthesis and reinforcement learning related work.
From UniPat's publicly disclosed projects in the past, it has been shown that the company is not content with simply ranking models.
In the field of software development, Monthly-SWEBench screens tasks from real GitHub projects every month to test whether agents can fix code, complete new functions and pass real verification. EvoCode-Bench allows models to continuously complete multiple rounds of tasks in the same workspace, requiring them to handle requirement changes, context accumulation and dependencies between previous and subsequent steps.
The Echo system extracts clues from news, data, prediction markets and real-time web information, generates probability judgments, and verifies the results after future events actually occur.
It does not test whether the model can answer a question whose answer is already known, but whether the model can make relatively reliable judgments about the future.
In a corporate press release issued by UniPat in April, the team let 5 EchoZ agents conduct a one-week real-market test on Polymarket, 4 of which obtained positive returns; the company also announced a 63.2% market win rate on political and governance issues, and a 59.3% win rate in predictions over 7 days.
This turns UniPat from an "AI examiner" into a more complex role: it not only tests models, but also trains models; it not only produces data, but also tries to turn prediction capabilities into sellable services.
What Alibaba values may be this layer of data infrastructure beyond the models themselves.
The most interesting part of Echo is that it turns the "future" into training data.
Traditional model training relies on past data. The idea adopted by Echo is "Train-on-Future", which means letting the model give probability predictions for events that have not yet occurred, and obtain result feedback after the events end.
UniPat's ExpertEval also adopts a similar idea. Its publicly released test set covers three high-risk fields of healthcare, finance and law, including 3213 expert cases, 207 scenarios, and an average of 21.83 scoring criteria per question.
These scoring criteria not only judge whether the answer is correct, but also judge whether the model has missed key conditions, violated professional rules, and made seemingly reasonable mistakes that have high real-world costs. UniPat refers to these mistakes as "critical negative items".
The judgment behind this method is that the bottleneck of model capability does not always lie in parameters and computing power, but may also lie in supervision quality.
UniPat stated in public materials that small models trained with expert Rubric supervision have achieved significant improvements in multiple professional tasks.
When computing power becomes expensive, how to improve the effective capability generated by each unit of computing power will be more important than simply increasing the training scale.
This is also the adaptive capability that Chinese AI companies may develop when facing computing power constraints.
For model training, these rules are new supervision signals.
The core assumption of UniPat is that high-quality supervision can improve the "unit computing power output" of the model. In other words, model capability does not only depend on the number of parameters and the amount of computation, but also on whether the training data can accurately point out why the model fails.
This direction exactly coincides with the problems the large model industry is facing: easily accessible Internet text is becoming insufficient, and what is truly scarce is data that can describe complex work, professional judgments and real-world feedback.
In essence, this line of thinking still uses higher-quality supervision to replace more computing power.
How to Maintain Third-Party Fairness, Independence and Autonomy
UniPat's development direction corresponds to the changes in Alibaba's AI strategy.
Alibaba is pushing AI from model R&D to cloud, enterprise software and consumer businesses.
In March 2026, Alibaba launched Wukong, an enterprise-level agent platform, hoping that multiple agents can enter enterprise workflows and gradually realize the transformation from dialogue to execution.
Alibaba has also established the Alibaba Token Hub business group, placing foundation models, model services and AI applications in a more closely connected organizational system. The official statement of Alibaba is to let AI complete the full chain from Token creation, delivery to application.
By June 2026, Alibaba merged the Tongyi large model team and the Future Life Lab into Token Foundry, which is directly managed by group CEO Wu Yongming. The outside world generally regards this as Alibaba's organizational move to further integrate model, application and commercialization capabilities.
What Alibaba needs is therefore not just a more powerful Qwen.
It also needs to know:
Can Qwen complete tasks in real office workflows? Are the answers generated by the model worthy of enterprise adoption? Can AI capabilities really be transformed into value that customers are willing to pay for?
All these problems require independent test environments, repeatable tasks and clear acceptance criteria.
Alibaba's performance data released in 2026 shows that the external revenue growth rate of its Cloud Intelligence Group reached 40%, and AI-related products have become an important growth driver for the cloud business; the number of Model Studio customers increased 8 times year-on-year. The company later disclosed that the revenue of AI cloud and computing power services reached $7.1 billion, a year-on-year increase of 45%.
The strategic logic of Alibaba's investment in UniPat may lie here: to purchase a set of judgment systems in advance for Qwen and more AI products in the future to verify whether they "can actually work".
If the financing is finally completed, UniPat's valuation of about $2.5 billion comes from three bets on future development.
First, whether UniPat can become a test supplier for model companies.
Second, whether it can master enough real tasks, professional rules and feedback data.
Third, whether it can make its test results enter the model procurement, enterprise deployment and model training processes.
Once the third point is realized, the evaluation company will no longer be an affiliated role in the model industry chain, but will become a kind of "standard setter".
But the risks are also clear.
Evaluation companies must maintain credibility. If they participate in the training of a certain model and are also responsible for evaluating that model, external customers will question their independence.
A bigger problem is whether the evaluation will be reverse optimized by model companies. As long as the test set, scoring criteria and task structure are fully studied, the model may learn to target the leaderboard, instead of learning to actually complete the work.
There will be more and more models, and more and more leaderboards. But what is really scarce may be a credible answer:
Can this model really complete the tasks assigned by customers?
Whoever masters this answer may stand in the acceptance link of AI commercialization.
What UniPat needs to do is to grow from a research-oriented evaluation company to the provider of this set of acceptance criteria.
This article is from WeChat official account "Digital Intelligence Surge", author: Zhou Yi, editor: Shen Xiao, published with authorization from 36Kr.