The world model has become the new hot track for embodied intelligence.
Entering 2026, the focus of the robotics industry is shifting from the robot body to the "brain", which has driven the World Model to gain rapid momentum.
Lyuyao, Chief Scientist of Light Elephant Lab, noted in a recent interview with media including Jiemian News that the evaluation of a World Model should not only focus on whether the next generated frame is realistic, but also check if the model can correctly explain why the world changes correspondingly when an action is adjusted. From his perspective, general embodied intelligence can only be achieved when the model masters the underlying causal laws of the physical world.
However, the industry still has no unified answer on what architecture robots should adopt to learn actions after they understand the world.
The discussions on the "brain" in the embodied intelligence field mainly revolve around two technical ideas: VLA and World Model. The two are not strictly mutually exclusive, but their core focuses are different:
The first path is VLA (Vision-Language-Action model), which connects inputs such as vision and language with robot actions. A typical VLA model learns to generate actions based on observations and instructions through robot demonstration data.
The other path is the World Model. Whether it is the generative model represented by Sora, the spatial intelligence model represented by World Labs, or the visual representation model represented by JEPA, they all essentially follow the same modeling paradigm: learning the statistical representation of the world from massive data to predict future changes from observations.
The debate over technical routes did not emerge this year. At the 2025 World Robot Conference, Wang Xingxing, founder of Unitree Robotics, publicly questioned the then-popular VLA route.
He argued that the biggest bottleneck for the large-scale application of humanoid robots currently lies in the AI model, and the root cause is the model architecture problem, rather than the hardware or data that the industry pays more attention to. He also described VLA as a "relatively dumb architecture". In contrast, he is more optimistic about driving robots through video generation models or World Models, believing that this route has a "higher convergence probability".
A year later, the World Model has obviously gained much higher popularity, but the technical routes have not converged as a result.
Regarding the VLA route, Lyuyao believes that it has hit the ceiling: the front end is a large language model, and the back end is connected to an action expert head. There is a modal gap when mapping language tokens to the action space, and it is almost unrealistic to generate truly generalized general policies by training VLA with only super-large-scale data.
Lyuyao said that no matter the industry turns to the World Model due to fierce VLA competition, or discovers the bottleneck of VLA in internal verification, the industry has begun to re-examine the original technical route.
However, the consensus basically stops here, and the concept of the World Model still lacks a unified definition so far.
The most intuitive manifestation of the lack of a unified definition is that the "World Model" mentioned by different people may be completely different things. Li Feifei and the World Labs team wrote a special article to discuss this concept this year, dividing the World Model into three categories according to functions: renderers, simulators and planners, and stated that "World Model" has become one of the most important and most easily misused terms in the current AI field.
According to this classification, the renderer mainly generates pixel images for human viewing, the simulator requires more accurate description of geometry, physics and dynamic changes, and the planner is responsible for determining the next action based on observations and goals. World Labs also classifies VLA as a type of planner, which means that VLA and the World Model are not necessarily mutually exclusive.
The differences are not only reflected in definitions. Guo Yandong, founder and CEO of Zhifang Technology, stated at the Beijing Academy of Artificial Intelligence Conference in June this year that the World Model is not a competing route of VLA, but a core component of the VLA system.
Beyond the division of technical routes, there are more practical problems. A person in charge of a World Model company told the reporter of Jiemian News that some video generation results look good at first glance, but their expression of physical properties such as friction is still insufficient, which is not enough for actual robot operation and execution. Some relevant data still needs to be collected by enterprises themselves at present, and the cost is not low.
The person in charge told the reporter of Jiemian News that the industry has not yet formed a unified standard for how to grade robot capabilities. At this stage, robots have been able to complete some processes with clear boundaries and short workflows in factories. But for open tasks such as household cleaning, robots not only need to identify the environment, but also understand the goals and causal relationships, and continuously adjust actions according to the execution results. There is still a big gap before they can truly complete such tasks.
Light Elephant chooses to start from the question of "how actions will change the world". Recently, Light Elephant Technology, in cooperation with the research group of Professor Li Shengbo from Tsinghua University, released the first generation of physics-native World Model Phi-WM 1.0 ActEffect, whose core is to transform the World Model into a "feedbackor during training", so that robots can learn "what consequences an action will bring".
Specifically, ActEffect can generate three action schemes with different levels of fineness at one time, predict the possible results of each scheme, and select the optimal scheme through comparison. After the training is completed, the World Model is removed from the inference link, and the policy network directly outputs the action through one forward calculation during deployment.
According to Lyuyao's explanation, it is equivalent to distilling the judgment of the World Model on the consequences of actions into the policy weights in advance, so as to reduce the inference burden in the deployment phase.
Apart from the model, data is also an unavoidable problem. The industry still has no unified answer on whether the World Model can reduce the data demand of robots.
Regarding the statement that "the World Model will reduce data demand", Lyuyao believes that it needs to be discussed on a case-by-case basis. He does not agree that the pure video prediction World Model can naturally reduce data demand: training a video prediction model with sufficient fidelity and generalization ability still requires massive data, and it is not necessarily reasonable to rely on it to generate data to fill the gap.
"But if the World Model can extract the underlying physical laws that are common across scenarios from the data, it may reduce the demand for data across scenarios and across robot bodies," Lyuyao said.
Zhang Tao, founder and CEO of Light Elephant Technology, gave an example to explain: if you throw ten objects of different weights and shapes into a basket, you can throw them in after ten practices, but you still need to re-try when facing the eleventh or hundredth unseen object. According to this logic, the scarcest resource is not the successful demonstration, but the data of a large number of trials and counterfactual reasoning in the simulation environment.
In this regard, Lyuyao judged that the required counterfactual trial and error data may be 10 times or even 100 times that of the real successful demonstration; but when the real machine post-training is carried out at a specific workstation, the data volume is about ten to dozens of hours, and no more than 100 hours at most. According to Light Elephant's idea, the World Model puts more trials and errors in the simulation environment with relatively lower cost, so as to reduce the demand for part of the expensive real machine data.
But factories will not only look at how much data the model needs. In the commercial landing stage, the debate over technical routes will eventually turn into a debate over ROI.
Research released by IDC in August this year shows that the market size of China's industrial embodied intelligent robots reached about 5.74 billion yuan in 2025. Its user survey shows that more than 80% of manufacturing enterprises hope that the industrial embodied intelligent robot projects can achieve investment recovery within two years.
This means that for industrial customers, it does not matter much whether to choose VLA or the World Model. What matters is whether the solution can be deployed quickly, run stably, and achieve payback within the expected time.
Although the technical routes have not yet converged, it has not prevented capital from pouring in rapidly. Data from CB Insights shows that the investment related to the World Model has increased from 1.4 billion US dollars in 2024 to 6.9 billion US dollars in 2025, nearly 5 times the original amount.
Entering 2026, large-scale financing is still ongoing. In February, World Labs founded by Li Feifei announced the completion of a new financing of 1 billion US dollars; in March, AMI Labs co-founded by Yann LeCun completed another financing of 1.03 billion US dollars. World Labs focuses on spatial intelligence and World Model, while AMI Labs takes modeling, reasoning and planning of the real world as its core direction.
Capital has placed bets in advance, but the speed at which robots really enter real scenarios is not that fast.
"The embodied intelligence industry is likely to experience a major reshuffle like the autonomous driving industry in the past," Lyuyao judged to media including Jiemian News. The capital market will eventually return to rationality, and enterprises lacking core technical moats and real delivery capabilities will be eliminated first.
This reshuffle may come soon. Zhang Tao believes that if the industry maintains the current state, a cold winter may indeed come within one or two years, but the result is not doomed. In his view, the real variable is whether the first embodied intelligent application verified by the market, with large-scale potential and capable of creating real commercial value can emerge before that.
This article is from "Jiemian News", author: Xu Meihui, published with authorization from 36Kr.