HomeArticle

Tsinghua's embodied model tops the global ranking, surpassing GPT-6 and Nvidia without relying on plug-ins or extra data.

量子位2026-10-09 11:48
From predictive generalization to action generalization

Whoa, on the World Action Model (WAM) track, a Chinese team has directly rushed to the global top spot?!

Moreover, a large number of world-leading embodied models such as GPT-6-Astra, π0.5 from Physical Intelligence, and NVIDIA GR00T-N1.7 have all been left far behind this time.

The player behind this achievement is a Chinese robotics company — the only direct affiliated embodied enterprise fully invested by Tsinghua University, Starbot Era.

Recently, VPP2, the self-developed world action model from Starbot Era, has successfully taken the first place on the RoboDojo simulation benchmark list.

It records a comprehensive average success rate of 32.26% and a comprehensive average score of 39.26 points, ranking first in both indicators.

The key point is that this competition is far from easy to win.

RoboDojo is known as the "Mount Everest" in the embodied intelligence industry, led by the MMLab of the University of Hong Kong, and jointly built by nearly 20 top global academic institutions.

It specifically sets questions based on the scenarios that robots are not good at: can the robot still work normally in a changed environment? Can it grasp the target accurately if the object is placed crookedly? Halfway through the task, can the robot still remember what happened earlier?

In short, it is not easy to muddle through by simply brushing up proficiency.

So what's the result?

Starbot Era's VPP2 not only won the overall first place, but also took the lead in the three dimensions of generalization ability, fine operation and memory ability.

It is reported that this time VPP2 did not add extra data, nor did it use enhancement methods such as Agent RSI. Relying only on the "standard dataset", it outperformed a large number of top competing models.

This means that the VPP2 model itself is strong enough, and its performance improvement comes from pre-training with a sufficiently generalized base model. Therefore, it does not rely on external plug-in enhancement strategies or expanded data to "brush scores".

Wait a minute, how on earth was such a powerful model trained?

We dug into its technical route, and it turns out that the breakthrough of VPP2 this time is really remarkable.

It targets a long-standing tough problem of world action models: if the prediction is wrong, the subsequent action will go wrong accordingly.

The solution proposed by VPP2 is to first train the video prediction ability to a sufficiently strong level, and then let the robot move in unfamiliar environments.

From prediction to action, both generalization capabilities are improved at the same time.

Judging from a number of currently disclosed test results, VPP2 has become one of the most competitive players on the WAM track of world action models.

World Action Model WAM, a new champion stands out from fierce competition

There is a rather awkward phenomenon in the embodied intelligence industry.

The robot demos are getting more and more dazzling, and the scores on the model list are getting higher and higher, but if you really ask which robot is smarter and more capable of doing practical work, there is no clear easy answer.

For example, when asking a robot to grasp a cup, it may hit the target 100% of the time on the desktop it has seen during training. However, once you change the cup or shift its position, the robot will immediately get confused and fail to complete the task.

Not to mention those tasks that require multiple consecutive steps to execute.

This is why the RoboDojo evaluation system is worthy of attention.

Its goal is to establish a unified and reproducible evaluation standard for embodied intelligence. The simulation test includes 42 dual-arm operation tasks, covering five dimensions: generalization, precise operation, long-horizon tasks, memory, and open-vocabulary instruction understanding.

To put it bluntly, it pulls robots out of their comfort zones.

Traditional robot models may get nearly full marks on familiar tasks, but once the environment, object position or task combination changes, their success rate may drop significantly.

RoboDojo specifically investigates such complex situations to test how much generalization ability robots have when facing changes.

The report card handed over by VPP2 this time is quite impressive.

With an average success rate of 32.26% and an average score of 39.26 points, both indicators rank first on the list.

For comparison, in the post-training evaluation of the RoboDojo simulation benchmark, GPT-6-Astra has an average success rate of 22.48% and an average score of 28.97 points.

VPP2 leads by 9.78 percentage points and 10.29 points respectively.

△ Screenshot from the paper, the baseline model data is taken from the official leaderboard

The two indicators have different focuses: the success rate measures whether the robot can complete the task completely; the average score measures how far the task has progressed, which can reflect the phased performance even if the task is not fully completed.

Both indicators integrate the test results of the five capability dimensions.

In other words, VPP2 not only has a higher proportion of completed tasks, but also shows stronger phased completion ability for tasks that are not finished.

Moreover, this lead is not limited to the overall performance.

Looking at the five capability dimensions separately, VPP2 also ranks first in the three dimensions of generalization, precise operation and memory.

These three capabilities correspond to several key thresholds for robots to enter the real world: whether they can work in a changed environment, whether their operations are accurate, and whether they can remember the previous steps when performing tasks continuously.

Leading in overall performance, and also competitive in key capabilities.

However, no matter how difficult RoboDojo is, the test environment is still a "simulated environment". When it comes to the physical world, variables such as object friction, deformation and position deviation will suddenly increase.

Can the generalization ability demonstrated by VPP2 in simulation still hold on real robots?

Starbot Era has also conducted relevant experiments.

This time, the team directly deployed VPP2 on a real ALOHA dual-arm robot, and tested 10 types of zero-shot operation tasks including grasping, placing, stacking, folding, and pouring.

In other words, without additional fine-tuning for these test tasks, the robot can directly perform the operations.

As a result, VPP2 once again delivered leading results: the average success rate is 58.5%, which is higher than the 40% of π0.5.

Among the 10 types of tasks, VPP2 won the best performance in 9 of them.

This is quite interesting.

The first place on the leaderboard proves its capability in the standard evaluation system; the zero-shot operation on real machines further examines whether these capabilities can be transferred to the real physical environment.

With the two report cards placed together, the technical advantages of VPP2 are even more worthy of in-depth exploration.

It is worth noting that VPP2 follows the World Action Model (WAM) route.

The basic idea of this type of model is to understand how the physical world will change next with the help of video prediction, and then convert this prediction into robot actions.

However, this route has long had a problem: a good-looking video prediction does not mean that the robot can really do practical work.

This time VPP2 is designed to solve exactly this problem.

From prediction generalization to action generalization, how does VPP2 achieve this?

To understand the breakthrough of VPP2, let's first look at a simple task: ask the robot to put the cup on the table into the box.

For humans, reaching out to grab, lift up, and put in requires almost no thinking.

But the robot has to judge the position of the cup, the motion trajectory of the robotic arm, when the gripper closes, and also predict what changes will take place in the physical world during the whole operation. Any deviation in any step may lead to failure.

The existing World Action Models (WAM) are exactly prone to problems at this point.

On the one hand, ordinary video models are good at generating images, but the images that look reasonable are not the same as those that conform to real physical laws.

For example, if the model is required to grasp the cup on the left, it predicts to grasp the cup on the right instead. The video can still be generated normally, but the robot will deviate directly when executing.

On the other hand, directly adding action learning to the video model may also damage the original generalization ability.

As a result, the robot learns the actions, but cannot perform them in a changed environment.

VPP2 puts forward a very straightforward judgment: the quality of video prediction determines the upper limit of actions.

Based on this idea, Starbot Era chooses to train video prediction and action learning in stages. The focus is not on "mixing video and action together", but to "decouple" the two and re"sort" their training order.

Starbot Era has designed a set of three-stage training strategy for this purpose:

In the first stage, event-level video continues pre-training, so that the model learns to predict the complete operation process.

In the second stage, fixed-duration video post-training and distillation are carried out, so that the prediction ability can keep up with the real-time execution speed of the robot.

In the third stage, action expert training is conducted to convert video prediction into specific actions, while preserving the original generalization ability as much as possible.

The three stages progress step by step, and finally lead to two core breakthroughs: prediction can be generalized, and action can also be generalized.

The first breakthrough: endow video prediction with generalization ability

The first problem VPP2 needs to solve is whether the robot can accurately predict the physical changes in unfamiliar tasks.

Based on the open-source Wan2.1-I2V-14B from Alibaba, Starbot Era integrates multiple types of data such as robot operations, human activities, and general videos, covering different robot bodies and operation modes.

However, the really interesting part is its data processing method.

The team did not simply feed all the videos to the model at once, but first divided the operation process into semantically complete segments, and then matched them with detailed descriptions.

For example, instead of just telling the model "put the cup into the box", it is necessary to specify which robotic arm, which cup, which box, and the entire operation process.

Because the same vague instruction may correspond to countless action trajectories.

If the model does not even figure out the operation object, it will naturally be difficult to accurately predict the future.

VPP2 describes the task in more detail, so that the language instruction and physical operation form a more stable corresponding relationship.

The more critical step is event-level video prediction training.

Traditional short-term prediction focuses on what will happen in the next few frames, while VPP2 directly lets the model learn the changes of a complete operation from the beginning to the end.

From the robotic arm approaching the cup, to grasping the cup, and then putting it into the box, the model needs to understand how the whole event happens.

In this way, when facing new objects, positions or even new operation tasks, it is more likely to predict reasonable physical changes.

In the instruction following test of robot operation videos, VPP2 with 14B parameters achieves a success rate of 90%, while the 64B parameter Cosmos3 model only reaches 78%.