HomeArticle

Alibaba plans to invest 300 million U.S. dollars to seize the "grading right": in the second half of the large model track, whoever defines "good" will determine the next generation of the industry.

Tao财经2026-09-28 17:14
"Examination Grading Right": Whoever defines what is good decides the next generation

Reports say Alibaba is considering leading a roughly $300 million investment in UniPat, an AI evaluation and post-training company, at a valuation of around $2.5 billion, with Tencent and existing shareholder HSG reportedly expressing interest in participating. It must be stated upfront that this transaction is still in the negotiation stage: the amount, valuation, terms and final investor list have not been finalized, it is far from being "closed", and everything may change at any time going forward.

But what is more worth discussing than the transaction itself is the core pain point that UniPat targets. Founded at the end of 2025, its founder Kuan Li is a post-90s generation entrepreneur. Before starting his business, he interned at Alibaba Tongyi Lab, participated in the Tongyi DeepResearch project, and focused on post-training analysis, data synthesis and reinforcement learning. The team has conducted complex reasoning research such as WebSailor, and released a series of evaluation benchmarks including BabyVision, SaaS-Bench, EvoCode-Bench and Terminal-X, with research directions covering software engineering agents, browser operation, and enterprise SaaS long-process tasks.

Simply labeling this company as an "AI examiner" underestimates its value. Its core position is the most easily overlooked layer in the second half of the large model industry: after the model outputs an answer, who makes the judgment, how to correct errors, and how to convert error correction into training signals for the next round. This layer is precisely the most scarce part of the entire industry at present, and the part that capital is most likely to misjudge.

After the model outputs an answer, who makes the judgment?

The first half of the large model industry competition centered on parameters, computing power and corpora. To judge whether a model is good or not, people relied on Benchmarks — input a question, output an answer, and calculate the accuracy rate. This set of methods was sufficient and simple enough in the "question answering" era.

However, it fails in the Agent era. Because the task of an Agent is no longer a single question, but a series of actions: understanding the environment, planning steps, invoking tools, modifying code, querying databases, handling exceptions, rolling back, and delivering results. There are dozens of steps in the process, and if any step goes wrong, no matter how perfect the final answer is, it is useless. Traditional static question sets cannot test "whether it can actually complete the task".

The first layer of UniPat's business is "production-level evaluation": instead of asking the model "whether it knows the answer", it puts the model in a near-real environment to test whether it can complete the actual task. SaaS-Bench loads 23 real open-source SaaS systems into Docker and runs 106 cross-application long-process tasks; EvoCode-Bench tests continuous development under multiple rounds of requirement changes; Terminal-X tests complex engineering tasks and version upgrades. This upgrades evaluation from "grading exam papers" to "actual combat drills".

Take the most intuitive example. Traditional Benchmarks test a math problem, and the model is considered good if it gives the correct answer. But in the Agent scenario, when the model receives the task of "help the customer modify a purchase order in the ERP system", it must first understand the requirement, then find the correct page, modify the correct field, save the change, and confirm that no other data is damaged. If any step breaks in the process, the task fails — even if the final output sentence it generates looks perfect. This is why "question-answering type evaluation" is completely insufficient in the Agent era: it tests "whether the model knows the answer", while what enterprises really care about is "whether the model can complete the actual task".

But the truly valuable part lies in the second layer.

After the evaluation is completed, the most valuable thing is not the score, but the failure trajectory: which step of decision-making went wrong, why the validator rejected the result, how the manual Rubric deducted points, and whether the task can converge if switching to another path. These trajectories, intermediate states, verification signals and preference annotations can be directly used for SFT, reward models or reinforcement learning.

Thus a closed loop is formed: identify weaknesses through evaluation → extract failure trajectories → label Rubrics and preferences → synthesize training data → update strategies via post-training → conduct re-evaluation → identify new weaknesses. What UniPat sells is not a one-off exam, but a feedback loop that enables the model to keep improving. The report also mentions that its cooperation with Alibaba focuses on reinforcement learning post-training data: it first conducts low-cost experiments on smaller models, screens out effective solutions and then scales them up. This "small first, large later" strategy is itself the most pragmatic path for the post-training data business — low-cost trial and error, and then large-scale deployment after the model is verified to work.

"The Right to Grade Papers": Whoever Defines "Good" Determines the Next Generation of Models

The term "the right to grade papers", when translated accurately, actually means mastering three core elements.

The first is the right to define tasks. What counts as "task completion"? Is it that the final JSON is correct, or that the cross-system state is consistent? Is it that the code can run, or that the test passes without affecting the existing logic? The way tasks are set will force the model to optimize in the corresponding direction. The party that sets the questions is actually determining the development direction of the model.

The second is the right to set scoring standards. Does it rely on rule validators, or Rubrics plus human experts? How to give process scores? Should safety, compliance, cost and rollback capability be included in the objective function? Once the standards change, the entire ranking of model performance will be rewritten.

The third is the right to feedback data. The most scarce resource is not "question-answer" pairs, but "environment-action-result-reward" data. Whoever can stably produce verifiable trajectories will control the supply end of post-training.

Therefore, "the right to grade papers" = evaluation standards + reward signals + source of training data. Its value does not lie in how authoritative a certain Benchmark is, but in the fact that the model can be replaced, while the feedback layer will be called repeatedly. Whoever masters the right to define "what counts as good" will participate in determining the development direction of the next generation of models.

These three elements are closely linked. Task definition determines what the model is trained to do, scoring standards determine what the model is trained to optimize, and feedback data determines what signals the model uses to make improvements. A company that masters the full set of these elements holds the answer to the ultimate question of "what counts as good". In the second half of the large model industry, this answer is far more valuable than "who has more parameters" — because parameters can be stacked and computing power can be purchased, but there is only one right to define "good".

The Rare Co-investment of Alibaba and Tencent, and Three Unavoidable Barriers

Alibaba's logic is the most straightforward. Tongyi needs post-training, and Alibaba Cloud's enterprise customers need acceptance testing. Enterprises will not pay for a model no matter how high it ranks on the public leaderboard; whether it can enter the ERP system to modify documents, enter the customer service system to close work orders, and enter the R&D workflow to complete unit tests determines the procurement decision. If capabilities like UniPat are embedded into Alibaba Cloud's delivery standards, Alibaba will hold the acceptance ruler for "whether the model works" in its own hands. Coupled with the trust chain of former Tongyi employees, the cost of technical alignment is very low.

If Tencent participates in the follow-on investment, its logic is different: it does not necessarily bet that a certain set of UniPat's Benchmarks will remain popular for a long time, but bets on the "model-neutral" feedback infrastructure. WeChat, WeCom, advertising and customer service, no matter whether they are connected to Hunyuan or external models in the future, all need to answer the question of "whether the Agent is stable and whether the process is smooth". Investing in the evaluation layer is more fundamental than betting on a single model — models will be replaced, but the "acceptance" work will always require a measuring ruler.

The rare co-investment of Alibaba and Tencent does not signal personal friendship, but a consensus: evaluation and post-training data are not minor supporting features, but preconditions for procurement in the Agent era.

However, outside the closed loop, there are three unavoidable barriers.

The First Barrier is Neutrality

With Alibaba leading the investment and Tencent participating in the follow-on, while its customers include other large model vendors, UniPat's credibility will naturally be questioned: will you relax the Rubric for your own model? Scale AI is a lesson from the past. UniPat can either make its validators completely transparent, or separate commercial evaluation from public Benchmarks, otherwise the "examiner" will sooner or later be questioned as a "private coach for its own team". This is an especially unavoidable barrier in the Chinese AI circle — because the customers and shareholders are exactly the same group of tech giants.

The Second Barrier is Revenue

What can be confirmed now is the technical reputation of its evaluation sets, while what cannot be confirmed is the sustainable contract and gross profit structure. The unit price of customized post-training data is high, but it relies on the budget of top AI labs — once the top labs cut their spending, the revenue will fluctuate sharply.

The Third Barrier is Engineering Capability

The maintenance cost of real SaaS environments, browser environments and multi-tool invocation scenarios is extremely high. It is normal that Benchmarks are constantly being broken by new models, and misjudgment by validators will directly contaminate the reward signals. Whether it can turn evaluation into an invokable RL environment, instead of becoming outdated right after the leaderboard is released, determines whether it will end up as a data company or an infrastructure company.

These three barriers are actually three sides of the same thing: UniPat wants to be a "neutral infrastructure provider", but one of its major shareholders is the model vendor Alibaba, its revenue relies on the budget of these model vendors, and its engineering system must keep up with the evolution of models at all times. A company that wants to be a neutral referee has to please the players, make a living from the players, and run faster than the players — this situation is destined to make every step it takes walk on a tightrope.

Back to the rumored $300 million investment. The first half of the large model industry competition focused on who has larger parameters, more GPUs and better corpora; the second half of the competition focuses on, after the model is released, how to confirm that it has actually improved, rather than only performing better on the leaderboard. Pre-training consumes the existing stock of internet data, while post-training consumes verifiable feedback. Once Agents are put into actual use, the correct answer is no longer a single-point output, but consistent state in long trajectories, successful tool invocation, exception recovery, and verifiable business acceptance.

Capital is targeting UniPat not for its current revenue, but for its potential to become the feedback nervous system for model companies: testing models, identifying weaknesses, generating data, giving rewards, and retraining models. Regardless of whether the transaction will be closed in the end, the direction has become clear: the players will continue to compete, and the examination rooms are starting to be contested. Once an examination room is occupied, the rules of the next competition will no longer be set by the players themselves.