HomeArticle

World models are becoming the "next crucial ticket to the future" for AI companies.

娱乐产业2026-08-18 16:46
Attempt to enable AI to truly integrate into the world

Over the past three years, large models have enabled artificial intelligence to learn to answer questions, generate images, and write code. However, when AI truly steps out of the screen and into vehicles, robots, factories and homes, being merely "capable of speaking" is far from enough.

A robot can accurately recite "how to put a cup into a cabinet", but it may drop the cup on the floor because it cannot judge the weight of the cup, the position of the cabinet door and the motion trajectory of the robotic arm; a car can identify vehicles ahead, but it may not be able to deduce the complex interactions between pedestrians, non-motor vehicles and other vehicles a few seconds later.

Language models learn the relationships between symbols, while what world models need to learn is: how the world works and what results actions will bring.

Entering 2026, world models have rapidly become one of the most crowded new tracks in China's AI industry. According to incomplete industry statistics, there are more than 30 relevant domestic enterprises. Technology companies such as Tencent, Alibaba, Huawei, Baidu have accelerated their layout, car companies including NIO have installed world models in mass-produced vehicles, and startups such as Jiazhi Shijie, Manifold Space, Daxiao Robotics, Xingdong Jiyuan, Tuoyuan Intelligence, Nijuzhen, LiberAI have received intensive financing.

Following large language models and video generation models, Chinese AI companies have begun to compete for the next ticket: enabling machines not only to describe the world, but also to understand the world, simulate the world, and act in the world.

Entering the physical world along four paths

The world model is not a fixed model architecture, nor does it have a unified industrial definition.

Simply put, it needs to build an internal "digital representation" of the external environment: understand spatial structures, object relationships and basic physical laws, predict future changes based on the current state, and evaluate the possible outcomes of different actions.

If large language models solve the problem of "what is the next word", world models try to answer: "What will happen next if I do this?"

At present, the exploration of Chinese companies is roughly carried out along four paths.

The first path is to move from video generation to an interactive world.

Video models can already generate realistic light and shadow, characters and motions, but a seemingly real video does not mean that the model truly understands the physical world. Characters may change positions out of thin air, objects may penetrate each other, and the same building may deform when viewed from another perspective. Based on video generation, world models need to add spatial consistency, long-term memory and motion response capabilities, upgrading the picture from "viewable" to "enterable and controllable".

The video generation capabilities accumulated by companies such as Kuaishou Keling, ByteDance Seedance, Shengshu Technology Vidu, Sand.ai are becoming important technical reserves in this direction. They have mastered massive video data and spatiotemporal generation technologies. Once the model can generate subsequent pictures in real time according to user actions, the video model may evolve into an implicit world simulator.

The second path starts from 3D generation and spatial reconstruction.

Tencent released and open sourced the Hunyuan 3D world model HunyuanWorld 1.0 in 2025, which can generate roamable, editable and exportable 3D spaces through text or images, and then launched Voyager that supports native 3D reconstruction and long-distance roaming. Game scenes that used to take modeling teams weeks to make are now being generated in minutes.

Companies such as Core Technology, VAST have accumulated experience in home decoration, architectural design and 3D content production. Compared with simply generating videos that "look real", this route puts more emphasis on geometric structures, spatial relationships and object attributes. Its commercial value is also more direct: game companies can quickly make maps, film and television companies can generate virtual scenes, and robots can be trained in advance in digital houses and digital factories.

The third path, which is the most concentrated direction for startups, is the combination of world models and embodied intelligence.

The core bottleneck of robots is not that they cannot understand instructions, but that they lack sufficient and high-quality physical interaction data. Language models can obtain trillions of levels of text from the Internet, but robots have to rely on real machine collection: a robotic arm grasping an object can only generate one trajectory, which is high cost, low speed, and has equipment loss and safety risks.

The world model is equivalent to building an infinitely replicable "virtual training ground" for robots. Robots can try repeatedly in the environment generated by the model first, and then migrate the learned strategies to the real machine, which can significantly reduce the cost of data collection and trial and error.

Jiazhi Shijie has a full-stack layout around the GigaWorld world model, GigaBrain embodied foundation model and robot body; Manifold Space has launched the WorldScape series, trying to directly support robot perception, prediction and action with the world model; Daxiao Robotics integrates understanding, generation, prediction and decision-making into a unified architecture; Xingdong Jiyuan proposes a world action model; Tuoyuan Intelligence chooses the Vision-World-Action route to allow the model to make predictions and plans in continuous physical space; LiberAI tries to combine human operation videos, UMI collected data and physical priors.

The "Wujie" series platforms such as RoboBrain and RoboOS released by the Institute for Artificial Intelligence, Beijing have also been open sourced, and have cooperated with many embodied intelligence enterprises to provide the industry with embodied brain and software-hardware collaboration frameworks.

The fourth path starts from intelligent driving.

Cars may be the world model carriers closest to mature commercialization at present. Car companies have large-scale real driving data, and the vehicles themselves are equipped with sensors, computing platforms and actuators, and safety, efficiency and user experience can all be converted into clear product value.

NIO released the NIO WorldModel in 2024, and the first version was pushed to more than 400,000 Banyan platform models in 2025 for scenarios such as active safety, high-speed navigation, urban navigation and parking. Baidu Apollo's ADFM, Huawei's world action model, and the end-to-end driving models of many car companies have also added prediction and simulation of future states to varying degrees.

This forms the basic layout of China's world model industry: video companies enter from pixels, 3D companies enter from space, robot companies enter from actions, and car companies enter from road scenes. Although the technical routes are different, the goals are highly consistent — to enable AI to obtain continuous cognition and action capabilities for the physical world.

Where will world models make money first?

The ultimate vision of world models is very grand, but their commercialization will not be achieved overnight.

In the short term, the first to generate revenue is not "general-purpose robots", but tools that can immediately reduce the costs of content production, data collection and testing.

The first type of commercial scenario is game, film and television, and virtual space production.

Traditional 3D content production requires multi-link collaboration such as concept art, modeling, material, lighting, and animation, which is high cost and long cycle. The world model can generate a complete space based on text, pictures or even a video, allowing creators to roam, modify and rearrange in it.

For game companies, it can be used for conceptual verification, map generation, level design and NPC training; for film and television companies, it can be used for virtual production, storyboard preview and digital scene construction; for e-commerce, cultural tourism and real estate enterprises, it can generate virtual exhibition halls, digital scenic spots and interactive model houses.

The business model at this stage is similar to the current generative AI, including model API calls, cloud computing power charging, software subscriptions and enterprise customization. Companies with game businesses such as Tencent can also verify the model in the internal production process first, and then gradually open it to developers.

The second type of scenario is autonomous driving simulation and safety testing.

It is difficult to recreate extreme scenarios on real roads: pedestrians suddenly crossing the road, blocked non-motor vehicles, and broken-down vehicles in heavy rain are all low-frequency but high-risk events. Relying on real vehicles to accumulate these data is not only a long cycle, but also may bring huge safety costs.

The world model can generate long-tail scenarios under different combinations of weather, roads and traffic participants in batches, allowing the driving system to undergo stress testing in the virtual world. Compared with charging consumers directly, this market is more suitable for charging by software platform, computing power usage, test mileage or project authorization.

The simulation platform may also become the infrastructure of the intelligent driving industry: one end connects the real data accumulated by car companies, and the other end connects model training, system verification and regulatory certification. Institutions such as China Automotive Engineering Research Institute have also incorporated large-scale synthetic data, world model deduction and long-tail scenario testing into the autonomous driving simulation system.

The third type of scenario is industrial robots, warehousing and commercial services.

Compared with homes, the environments of factories and warehouses are more structured, the task boundaries are clearer, and the return on investment is easier to calculate. After robots complete sorting, loading and unloading, handling, inspection and assembly, enterprises can directly measure how much labor is saved and how much efficiency is improved.

Therefore, world model companies are likely to enter industrial customers in the form of "model + robot + scenario solution" first, rather than selling a basic model alone. Model companies need to deeply integrate with robot bodies, sensors, industrial control systems and production processes, and continuously optimize through on-site data.

This is also the reason why companies such as Jiazhi Shijie, Manifold Space and Daxiao Robotics are developing towards full stack: without the body and customer scenarios, it is difficult to form a data closed loop; without a continuously evolving model, robot hardware products are easily trapped in low-price manufacturing competition.

The fourth type of scenario is home robots, but it may mature last.

The home environment seems simple, but it is actually more complex than factories. The house type, furniture, items and living habits of each family are different, and the elderly, children and pets bring a lot of unpredictable factors. Robots not only need to understand tasks, but also judge human intentions, and safely complete long-term operations in unfamiliar environments.

This means that what home robots need is not a model that can demonstrate folding clothes, but a system that can quickly recognize new houses, predict action results, handle failures and continue learning. Although the commercial space is huge, before the reliability and cost cross the critical point, the market will still mainly focus on high-end products, scientific research platforms and specific functional devices.

Building a data flywheel for the physical world

World models are repeating the early boom of large language models: the number of companies is increasing rapidly, the conceptual boundary is expanding, and financing and valuation are growing ahead of revenue.

However, world models cannot be simply understood as "more advanced video generation models". The more realistic a video looks, the more it does not mean that it understands physical laws. A truly commercializable world model must at least cross four thresholds.

The first threshold is physical consistency.

The model should not only generate the picture of a cup falling to the ground, but also understand why the cup falls; it should not only reproduce the car turning, but also know the constraints between speed, road curvature and tire state. Otherwise, the more data it generates, the more serious the wrong training it may bring to robots and cars.

The second threshold is long-term consistency.

Current models can maintain scene stability in a short period of time, but once the task is extended, object positions, spatial structures and causal relationships are prone to drift. Robots tidying up rooms and cars passing through complex urban areas all require stable memory and prediction that last for several minutes or even longer.

The third threshold is real-time performance and cost.

If it takes dozens of seconds for the model to deduce one step, it can produce videos, but it cannot control high-speed driving cars and continuously moving robotic arms. The world model must finally perform low-latency reasoning under limited computing power and power consumption, which will test the model architecture, chips and software-hardware collaboration capabilities at the same time.

The fourth threshold is the closed-loop verification of the real world.

In the future, the criteria for measuring world models should not only be the beauty of generated videos or laboratory rankings, but also the success rate of robot tasks, the autonomous driving takeover rate, the consistency between simulation and reality, and the safety and economy after deployment. A large number of "world first" evaluation results still need to be verified in more bodies, more environments and longer periods of time.

China has unique advantages in this round of competition. It has a huge manufacturing, intelligent vehicle and robot supply chain, as well as dense scenarios such as factories, warehouses, residences, roads and commercial services. What world models need is not simply Internet text, but real physical data from sensors, machines and production processes.

But scenarios will not automatically become barriers. Only companies that can enter customer sites, complete data collection, model training, simulation verification, equipment deployment and feedback update can transform industrial resources into a real data flywheel.

In the future, there may be three types of long-term players in the world model industry: one type masters general foundation models and cloud computing platforms, and exports APIs and development tools to the outside; one type deeply cultivates vertical fields such as automobiles, robots and industry, and realizes monetization through industry models and solutions; the other type masters 3D assets, physical data, simulation evaluation or ontology equipment, and becomes an infrastructure provider in the industrial chain.

On the contrary, companies that only have model demonstrations and lack real data and delivery capabilities are likely to be the first to be eliminated after the financing boom recedes.

Large language models allow AI to learn to talk about the world, while world models try to allow AI to truly enter the world.

It may not appear in front of ordinary people in the form of an independent app, but it may be hidden behind every car, every robot, every digital factory and every virtual space. The real commercial value of world models does not lie in generating a world as realistic as possible, but in allowing machines to "imagine" the future internally before taking action.

When this kind of imagination is accurate enough, fast enough and cheap enough, artificial intelligence can be said to have truly stepped out of the screen.

This article is from the WeChat Official Account "Entertainment Industry" (ID: yulechanye), written by Large Model, authorized for release by 36Kr.