Li Fei-Fei's team has developed a World Action Model: by switching to a new base, the success rate of robots has skyrocketed directly.
On September 22, Black Forest Labs, renowned for its FLUX image models, made a cross-domain move to release FLUX 3 Action, an open-source 7B world action model. According to its official blog, this model outperforms Cosmos 3 Nano, the previous strongest open-source model, on NVIDIA's RoboLab-120 benchmark, yet its parameter count is less than half of the latter. The fact that an image generation company is stepping into robot control R&D indicates that the "physical intuition" accumulated in video models is being treated as the foundation of robot brains by an increasing number of teams.
This technical path is the so-called World-Action Model (WAM). Since NVIDIA explicitly proposed this term in its DreamZero paper in February this year, a large number of relevant papers have emerged rapidly. However, behind the boom lies an awkward problem: almost all these systems have simultaneously replaced their video backbone, training data, video-action interaction mode and inference workflow, and no one can clearly tell which specific design brings the performance improvement.
This week, the team led by Li Fei-Fei, Wu Jiajun, Ehsan Adeli and other researchers from Stanford University published a paper, proposing an open world action modeling framework called OpenWAM, which aims to give a systematic answer to this problem.
Overview of the OpenWAM framework: three-stage training, shared MoT architecture, four configurable interaction modes, and a reusable local dynamics model that can be frozen for migration.
- Paper Title: OpenWAM: An Open Framework for Composable World-Action Models
- Paper Link: https://arxiv.org/abs/2610.07922
- Project Homepage: https://openwam.stanford.edu
- Code Repository: https://github.com/OpenWAM/OpenWAM
Why We Need a "Unified Test Standard"
The traditional Vision-Language-Action Model (VLA) is like a driver who drives entirely by conditioned reflexes: it directly outputs actions as soon as it sees the scene and receives the instruction.
WAM adds an extra step of "mental rehearsal": it simultaneously predicts what the scene will look like in the next few seconds, and what actions should be taken accordingly. Videos naturally record how objects fall, collide and are pushed, so models pre-trained on massive volumes of videos inherently carry physical priors.
The problem is that there are numerous ways to couple "predicting the scene" and "generating actions": imagine first then act, act first then imagine the consequences, generate both simultaneously, or compute them separately. Different teams choose different architectures, base models and datasets. Comparing their performances together is like several chefs changing ingredients, cooking stoves and cooking sequences at the same time, making it almost impossible to figure out which step actually makes the dish taste good.
The second problem is more hidden. Most robot training data consists of successful demonstrations, which only tell the model "how to take the correct path" and barely contain information about "what will happen if you stretch your hand two centimeters further". This is similar to a trainee who only watches the coach's perfect slalom videos: the trainee remembers the route, but does not know how much the car will deviate when the steering wheel is turned by a certain angle.
OpenWAM solves the first problem with a shared base model plus switchable interaction modes, and addresses the second problem with large-scale "counterfactual" data.
Same Backbone, Four Modes of "Think First or Act First"
OpenWAM is built on Wan2.2-5B, an open-source video model from Alibaba. It is further pre-trained on about 3.34 million videos of robot-human interactions, with a nominal total duration of around 14,600 hours. No action labels are used during this process, and the training runs on 32 B200 GPUs for 14 days.
The key modification is "causalization": each video block can only access content from the current moment and earlier, and cannot peek at future content. The original model is like someone who holds the full script and revises lines repeatedly, while the new model has to perform in chronological order like a live broadcast, which exactly matches the requirement for robots to act step by step.
Subsequently, the team combined the 5B video expert and 2B action expert using the Mixture-of-Transformers (MoT) architecture. Both experts retain their own normalization, projection and feed-forward layers, and only exchange information in one joint attention layer.
Shared MoT architecture: the video expert and action expert perform joint attention on the packed tokens.
Based on this backbone, the team defined four "interaction procedures":
- VTA generates actions after imagining the scene first
- ATV follows the reverse order
- Joint makes the two generate simultaneously and refer to each other
- Decoupled makes the two invisible to each other in their future segments
This is like the same band playing the same piece of music, with the only variables being who starts playing first and whether the members can hear each other. The four modes share the same base model, tokenization and flow matching objective, the only differences are generation order and cross-modal attention, which makes the comparison a typical controlled variable experiment.
Attention masks for the four interaction procedures (A–D) and two local dynamics procedures (E–F).
Separate and Train the "Action Translator" Independently
The more innovative part of the paper is splitting the Inverse Dynamics Model (IDM) and Forward Dynamics Model (FDM) into components that can be trained independently.
The IDM works like a translator: it takes the current scene, the robot's own state and a segment of imagined future video as input, and outputs specific actions. The FDM works the other way around: it predicts what the scene will look like according to the given actions.
The key point is the "local" property: neither of them receives task instructions or earlier historical information, they only focus on "what action corresponds to this small segment of scene change". Therefore, they are not bound to specific tasks, and can be spliced with any compatible video predictor. The video model decides "what should happen", while the IDM is responsible for "how to achieve it".
However, the translator that only learns from successful demonstrations has too limited exposure. To solve this problem, the team built the LIBERO-Long-CF dataset: they restore the simulator state in the demonstrations, execute modified actions such as stopping, reversing, adding noise, and changing the gripper opening/closing timing, etc. Finally, they obtained 32,000 segments with about 4.1 million control records, which is 29.7 times the size of the original demonstration dataset, and 72.3% of the segments changed the placement state of objects. Many trajectories cannot complete the task at all, but this is exactly the core value. Going back to the driving school analogy, this is equivalent to the coach asking the trainee to turn the steering wheel half a circle more on purpose, so that the trainee no longer only remembers the route, but also understands the actual handling characteristics of the car.
Experiments: The Base Model Performs Excellently, and the Translator Can Be Migrated
On the four subsets of LIBERO, all four interaction procedures achieve very high success rates. The average success rate of VTA reaches 98.6%, slightly higher than the previously reported results of LingBot-VA (98.5%) and π0.5 (96.9%) cited in the paper. However, the performance on LIBERO is almost saturated, and the more noteworthy detail is that on LIBERO-Long, the Decoupled mode with mutual invisibility in future segments (97.0%) performs on par with the bidirectional interactive Joint mode (96.6%).
Closed-loop success rates (%) on the four LIBERO subsets, with baselines being previously reported results.
On real robots, the team tested three tasks on two Franka FR3 robotic arms: toasting bread, restoring the last layer of a 2×2 Rubik's Cube, and sorting cups by color. The average success rates of VTA and Joint are 92.1% and 91.9% respectively.
Left: three real-world dual-arm tasks, Right: four LIBERO-90 tasks for component migration.
Ablation experiments show the importance of the base model: with all other conditions unchanged, the VTA initialized with the original Wan2.2 only achieves a 68.4% success rate on LIBERO-Long, which rises to 97.8% after replacing it with the causal base model pre-trained on robot videos, bringing a 29.4 percentage point improvement. The authors emphasize that this is the combined effect of data and causalization modification.
Impact of video backbone initialization method on the success rate on LIBERO-Long.
The most critical part is the component migration experiment. The team froze the IDM trained on LIBERO-Long, and paired it with video predictors adjusted for four new LIBERO-90 tasks.
The results show that the local IDM trained with both demonstrations and counterfactual data achieves an average success rate of 84.0%, while the IDM that receives full context only reaches 47.0%, and the local IDM trained only with demonstrations is as low as 21.5%.
"Only focusing on local segments" and "having seen a sufficient variety of consequences" are both indispensable. Only by combining the two can we get a migratable action translator. It is worth noting that the video predictor for target tasks is fine-tuned with demonstrations containing action labels, which is not zero-shot migration, and the authors have explicitly stated this point.
Success rates of migrating the frozen IDM to the four LIBERO-90 tasks.
The conclusion for FDM is consistent. Starting from the same state, 16 different actions are executed, and the model is required to judge which real result the predicted scene corresponds to. The random guess baseline is about 6%, the FDM trained only with demonstrations reaches 21.1%, and the value rises to 71.3% after adding counterfactual data, which proves that the model can finally distinguish the difference between "push this way" and "push that way".
Prediction performance of the local FDM on counterfactual transfer.
Conclusion
The authors also openly acknowledge the limitations: component reuse is mainly verified in simulation, and the cost of "making mistakes on purpose" on real robots is much higher; the FDM only covers short-term prediction. But the value of OpenWAM does not lie in improving the benchmark score by another fraction of a percent. At a time when WAM papers are emerging on a weekly basis, a public testbed with a fixed base model and only one variable changed can help the field shift from "who gets a higher score" to "why the score is higher". The reusable action translator also implies that future robot brains may not need to be black boxes trained as a whole: the module responsible for imagination and the module responsible for execution can be iterated separately and assembled on demand. Those trajectories that fail to complete the task are exactly the nutrients that help the model understand the physical world.
At present, the team has open-sourced the training, evaluation and combination code, as well as the pre-trained video model weights.
This article is from the WeChat official account "Jiqizhixin" (ID: almosthuman2014), author: AI-focused team of Jiqizhixin, 36Kr publishes this content with authorization