The Fierce Battle Over World Models: 55 Companies, 7 Camps, Who Is Actually Building Real World Models?
The embodied intelligence industry has been hit by another bombshell.
Shao Tianlan, founder of Mech-Mind, posted a fiery rant on WeChat Moments: a number of well-known, high-valuation companies that have even appeared on the Spring Festival Gala in Beijing and Shanghai are extensively using so-called "data collection centers" and related-party transactions with local governments, investors and suppliers to generate fake, unsustainable revenue. He pointed out that this practice is illegal, unethical and unwise.
Galbot later responded: The company's products have been deployed on a large scale in multiple sectors including industry, retail and pharmacy, and have completed two years of normalized operation verification.
On the surface, the dispute is over whether the robot companies' revenue figures are real, whether the data from their physical units is real, and whether their commercial scenarios are real.
But looking deeper, there is a bigger question behind it: is the "real world" that everyone claims to own actually real?
Coincidentally, if you ask 100 investors what areas they are focusing on this year, they will all tell you it is the world model, which is expected to enable machines to understand, predict and simulate the real world.
The direction sounds clear, but its boundaries are far from well-defined.
When we really ask the question what exactly counts as a world model?, 100 investors may not have 100 different answers, but their opinions have already started to diverge.
Apart from the players that have long been developing embodied brains, companies focused on video generation, 3D generation, data services and simulation platforms — enterprises that did not originally operate under the "world model" banner — have all begun to be included in this new narrative.
The world model has become a standard feature for every related or unrelated company, appearing in the headers of PPTs, the titles of press conferences, and the gilded sections of financing stories.
What is more interesting is that this track, which does not even have a unified definition of industry boundaries, has already started to churn out unicorns in batches: among the 55 companies we have sorted out, 12 have already reached unicorn status, 9 of which are from the embodied intelligence sector.
Market enthusiasm for this term is also growing far faster than actual product implementation. According to Google Trends data, the global search popularity for "world model" has skyrocketed 1200% year on year, surging rapidly from May to July 2026.
Google Trends: Global search popularity of the world model
The "UnDefined" official account recently sorted out 59 domestic world models and divided them into seven schools: the Robot School, the Autonomous Driving School, the Big Tech School, the Visual Generation School, the Data School, the 3D School, and the Native School.
Different schools, with different origins and different core advantages, are destined to take different paths to survival.
Next, we will give unfiltered reviews of these "world models" from the ground up, focusing on three key points: how close it is to a real world model, whether it has truly entered the real world, and whether it can continuously obtain feedback from the real world.
Of course, we are not a professional institution! All evaluations are not objective at all! Any resemblance to actual facts is purely coincidental.
01 The Robot School: Competing for the brain that understands the world — I rate it top tier
The robot school's development of world models is not only extremely crowded, but also driven by genuine rigid demand.
At present, there are at least 18 robotics and embodied intelligence companies that have publicly announced their world models: SenseTime Jueying (Kairos), Daxiao Robotics (Kairos 3.1), StarGis (Fast-WAM), Agibot (τ0-WM, GE-Sim 2.0), Galbot (DyWA, LDA-1B), Lingchu Intelligence (Psi-W0), Xingyuanzhi (ω-EVA), GigaVision (GigaWorld-Policy), Moqi Intelligence, Octopus Dynamics, Zhongke Fifth Era (BridgeV2W), Weitejia Warehouse Robotics (Ulti-Brain Embodied Brain), Zhijian Dynamics (VLA LaST₀), Zhi Pingfang (Full-domain Full-body VLA GOVLA), Feijie Cosmos, Wujie Dynamics (W‑WALA), Zili Robotics (WALL‑WM), Pok Robotics (Embodied World Model), etc.
The robot school has the highest capital density in the entire world model track. In the first half of 2026, the disclosed financing amount in China's embodied intelligence sector alone exceeded 460 billion yuan, 70% of which went to the top 20 companies. The valuations of Galbot, StarGis, Zhi Pingfang and Zili Robotics have all exceeded 20 billion yuan. Wujie Dynamics, which was established less than a year ago, has also seen its valuation exceed 10 billion yuan after four consecutive rounds of financing, while Zhijian Dynamics joined the unicorn rank in just half a year.
What capital is betting on is who can monopolize the physical unit data entry first.
Traditional robots rely on rule-based programming, and end-to-end models have very poor generalization ability. When encountering unknown objects, they will crash directly; when switched to a new scenario, they need to be retrained. This is exactly the problem that people expect world models to solve: after understanding physical laws, they can deduce consequences in their "mind" and choose the optimal action.
The biggest advantage for the robot school to enter the world model track is that they own real application scenarios and physical unit data (emphasis: real data that can be used to train capabilities, not the fictional data from the so-called data collection centers mentioned earlier).
Every missed grasp, collision, fall and human intervention of robots in the real environment is the most precious training data for world models. Once the scale of physical units grows, the flywheel of "model training - real machine verification - product iteration" will have the opportunity to operate continuously.
Therefore, the robot school is very likely to be the first group to bring world models into the physical world.
But the problem also lies here: physical units are assets, but also liabilities. Without enough physical units, the closed loop is just empty talk; but once the number of robots increases, the cost of hardware, deployment and maintenance may drag the company down first.
There is also a more hidden trap: a dedicated model is not equal to a general world model. The data generated by the company's own robots every day can indeed make the model more and more familiar with its own robotic arms, grasping methods and scenarios, but this may only make a "small world" more and more proficient. The real test is whether the model can still work effectively when switched to another robot, another task, or an unfamiliar environment.
Therefore, the final competition of the robot school is about who can turn physical unit data into transferable physical world knowledge, and evolve from "only understanding its own robots" to "understanding other people's robots". This is the real development of a brain that understands the world; otherwise, the world model is just a more advanced simulator for its own robots.
Only one or two players can achieve this transition; the rest are just using the name of world model to add gimmicks to their financing stories.
02 The Autonomous Driving School: Predict the world in the next second — I rate it top of the top tier
At present, at least 7 car companies and autonomous driving companies have publicly announced related explorations, including NIO (NIO WorldModel), Li Auto (Reconstructed World Model, Generated World Model), XPeng (X-World), Geely (World Action Model), Xiaomi Auto (Joint World Model), Huawei Cloud (world model based on the Pangu multimodal large model), Momenta (R7 reinforcement learning world model), etc. In addition, Horizon and Star River Wen Tu have also publicly stated that their self-developed autonomous driving systems have built-in world model capabilities.
The list is not long, but each player has its own mass-produced vehicle fleet and data pipeline.
Today's "autonomous driving" is essentially still assisted driving. What the world model needs to supplement is the unification of perception, prediction and planning.
The biggest advantage of car companies is that they have continuously running real-world samplers, which realize data closed loop through collecting real road data, discovering problems, screening effective samples, retraining, verifying model iteration effects through simulation and evaluation, and finally redeploying the model to vehicles.
The data collection conditions of the autonomous driving school are highly standardized: relatively fixed sensors and vehicle control interfaces; relatively consistent road conditions; relatively repetitive tasks, so the model has strong scalability for mass replication.
Different from other schools, the autonomous driving school has already completed a round of data closed loop construction in the past ten years, so they may become one of the earliest vertical fields to run through the world model.
But car companies also face very realistic problems: building cars and building foundational models are two different sets of capabilities. Having money, vehicles and data does not mean owning a ten-thousand-GPU cluster and a pre-training team. Therefore, the world models of some car companies may be "branded" or "jointly developed" — the name belongs to the car company, but the underlying capabilities come from suppliers.
This has also become the real hidden competition of the autonomous driving school. Suppliers such as Huawei and Momenta can obtain data across brands and vehicle models, especially covering high-value scenarios such as extreme weather, complex intersections and dangerous decision-making. More importantly, they have the opportunity to connect data, algorithms and evaluation systems of different fleets. Their ambition is to evolve from "selling autonomous driving solutions" to "providing foundational capabilities for the intelligent automotive era".
Similarly, suppliers also have their own troubles: having access to data does not mean owning the data closed loop. Whether the data can be continuously used for training ultimately depends on the cooperation mode and contract terms with car companies.
Car companies are competing for their own data scale and closed loop, while suppliers are competing for cross-fleet data scale and closed loop. In the end, whoever can truly turn data into model capabilities will be qualified to define the brain of next-generation vehicles.
03 The Big Tech School: Everything can be done, but the world cannot be built by piling up resources — I rate it top tier
At present, big tech companies that have publicly launched related products or research include: Alibaba (Qwen-RobotWorld, Qwen-AgentWorld, Happy Oyster), Ant Group LingBot (LingBot-World / LingBot-VA), Tencent Hunyuan (HY-World 2.0), ByteDance Seed (GR series, Seedance 2.0), Kuaishou (Keling Video 3.0 Omni), Huawei Cloud (Pangu World Model), as well as Meituan (LongCat-Video) that is currently exploring this direction.
In 2026, world models are increasingly regarded by big tech companies as the next generation of AI infrastructure. They can tolerate that their large models are not the strongest, but they can never allow themselves to be absent from this new game.
The biggest advantage of the big tech school is that they have all kinds of resources. Companies with video resources develop video-related models, companies with map resources develop 3D-related models, companies with robots develop embodied models, companies with Agent resources develop digital environment-related models... Alibaba even bets on three tracks at the same time: robots, Agents and virtual worlds.
This "bet on everything" confidence is an advantage in the period of ambiguous strategy, but once the standard is established, it may also become a burden. After all, there is still no answer on whether the world model should be a video, a 3D model, a physical simulator, or a digital environment for Agents.
You can bet on multiple routes at the same time, but you cannot make multiple data closed loops work at the same time. Content platforms have massive amounts of videos, but lack action feedback; map companies have spatial data, but may not have continuous physical interaction; robotics companies have real action data, but the scale is far smaller than internet data.
Having a large amount of data does not mean having the data most needed by the world model. What they finally produce may be a bunch of "advanced tools" that serve video, 3D, robot and Agent respectively.
The final outcome of the world model may not be won by the big tech company with the most models, but by the one that can connect models, scenarios, data and feedback into a closed loop, and run out a real data flywheel from all the scattered bets.
04 The Visual Generation School: From generating images to simulating the world — I rated it NPC level at the beginning of the year, now I rate it top of the top tier
For the visual generation school, developing world models is essentially re-pricing their existing capabilities.
Current companies with publicly related models include: Sand.ai (Magi-1.1), Zhixiang Future (native full-modal world model architecture), Moxin Tech (KOKONI-World / MoWorld), Shengshu Tech (Motus / MotuBrain), Alibaba Happy Oyster, as well as Kuaishou Keling 3.0 Omni and ByteDance Seedance 2.0 in a broad statistical sense.
This group holds the cheapest massive amount of internet videos — which are essentially slices of the physical world in the time dimension.
Sand.ai uses autoregression to predict the next video clip, Moxin Tech cuts in from the 3D space dimension, and Shengshu Tech integrates VLA, world model and video generation into Motus. From "generating the next frame" to "predicting the next state", the technical stack is inherently continuous.
Capital also recognizes this logic. Shengshu Tech's Series B financing is nearly 2 billion yuan, with a post-investment valuation of over 2 billion US dollars, and it is sprinting for a Hong Kong IPO; Zhixiang Future has completed three consecutive rounds of financing in three months, totaling more than 2.1 billion yuan; Keling AI completed independent financing of nearly 3 billion US dollars, with a pre-investment valuation of 15 billion US dollars.
What capital is buying is no longer just the "generating images" business, but also the imagination of evolving from content generation to physical simulation. Being able to generate a video of a cup falling on the ground sells as a content tool; being able to predict the trajectory of the fragments splashing after the cup falls sells as a physical understanding capability.
From "content generation" to "physical simulation", the valuation logic differs by an order of magnitude. The problem is that the premise of re-pricing is that the capability is really worth the price.
The core advantage of the visual generation school is that they have data, models and engineering capabilities, but their biggest shortcoming is that they have images, no causality.
Existing research also shows that visual generation models are still stuck in problems such as spatial reasoning, persistent state, long-term consistency and causal understanding, which are far from real world models. This is a paradigm problem — from "unconsciously imitating pixels" to "consciously understanding physics", there is a closed loop of "action-feedback" in between.
Therefore, what this school needs to cross in the end is the transition from visual continuity to physical causality, from "generating the world" to "simulating the world".
05 The Data School: From reconstructing the world to generating the world — I rate it top tier now, and I will adjust the rating later this year
Current companies that have publicly developed self-owned models include: Dayan Tech (4D world reconstruction synthetic data platform), Zhuoyin Intelligence (Simulaix), Shuwei Intelligence (embodied intelligence physical causal data engine), Qunhe Tech (SpatialGen), 51WORLD (51World Model) and Cross Dimension Intelligence (DexWorldModel).
First of all, in the embodied intelligence industry full of non-consensus views, data shortage has become one of the few consensuses at present.
Therefore, the data school, which is stuck at the bottleneck of the industry, has natural advantages. But the question is, after the data problem is solved, will they still be valuable?
In 2026, the world model is obviously a more attractive capital story than data labeling and simulation services. Therefore, it has become a consensus for such companies to launch their own world models one after another, and shift from charging by order to charging by capability.
The