Can the combination of AI videos and world models thoroughly catalyze the full maturity of interactive film games?
When the player presses the direction key, the road on the screen extends immediately; when the audience speaks, the virtual character not only answers the question, but also turns around, moves, and even changes the surrounding space — in the past year, such scenes have frequently appeared in the demo videos of world models.
Nowadays, ByteDance is also reported to have joined this competition.
According to Bloomberg, Zhang Yiming is personally overseeing a real-time spatial video model developed based on Seedance. Its goal is no longer to generate a fixed video, but to continuously generate an interactive dynamic environment based on user input.
Games, short dramas, live streaming and PICO may all become the future application scenarios of this technology.
ByteDance has not given an official response, but in the same period, Vidu S2 released by Shengshu Technology added another puzzle to this path: characters can respond to users in real time during video generation, and characters, costumes and backgrounds can also change instantly. Its spatial video version also plans to generate binocular images with depth of field for VR devices.
Overview of Vidu S2
From a technical perspective, AI video is trying to move beyond "generating a piece of content" and towards "generating a space that can respond to users".
However, from technical demonstrations to breakthrough progress in truly interactive film and games, there is still a whole set of narrative, product and commercial problems in between. Especially when observed on a global scale, this revolution has probably just stepped out of the laboratory.
From a single video to a responsive space
The technology collectively referred to as "world model" today actually includes many routes with great differences.
Google's Genie 3 can generate explorable environments based on text descriptions. Users can control character movement, and the world will continue to generate along with the actions. However, Project Genie open to users is still a research prototype, and the single experience time is only about one minute.
Runway's GWM Worlds 2 is already capable of continuously generating interactive videos with sound, and tries to cover preset films, turn-based interactive films and real-time games. But Runway also admits that pre-planned prompts are still more stable, and the model's ability to maintain long-term memory, spatial structure and object relationships is still limited.
Vidu S2 from Shengshu Technology demonstrates another path that is closer to entertainment applications.
It is in no hurry to generate an entire open world, but first allows characters to perform, respond and change in real time in the video. Its VR spatial video capability is particularly noteworthy: if the left-eye and right-eye images can be generated synchronously, AI video will no longer be just images pasted on the screen, but may become an immersive content with a sense of space.
However, this function is still in the preview stage. The currently public products are mainly virtual humans and real-time video editing. It is obviously too early to directly describe it as a mature VR world.
On the side closer to games, Odyssey and Decart are also exploring controllable generative environments; World Labs emphasizes sustainably editable three-dimensional space; some research institutions apply world models to robots and autonomous driving, allowing AI to understand objects, actions and their consequences.
Atlas from World Labs
Capital has poured in rapidly.
World Labs is valued at billions of dollars; Odyssey and Decart have also received large amounts of financing one after another... However, the problems faced by all companies at present are very similar: the generation speed is increasing, but the consistency of the world has not been solved synchronously.
After a door is opened, turning back may lead to a completely different room; the character can answer immediately, but may not remember what happened ten minutes ago; the model can continuously generate images for several minutes, but cannot support the plots, rules and causal relationships required for a two-hour movie or a game of dozens of hours.
Therefore, today's world models can indeed create the first impression of "being in the world", but cannot stably maintain a real world yet.
This is also the most confusing point for many domestic concepts of "AI interactive film and games".
AI interactive film and game work "Under the Jingmen Gate"
AI-generated art assets, free-talking NPCs, branching plots, real-time video editing and world model generated environments originally belong to different levels of capabilities. Packaging them in the same product does not mean that the boundary between interactive films and games has been broken through.
A chatty character does not equal having memory and motivation; a modifiable image does not equal that the user has the right to act freely.
After generation, who is responsible for telling the story?
This is also the most easily overlooked issue when discussing real-time interactive content at present: technical demonstrations are rapidly becoming amazing, but products that can be consumed by users for a long time have not emerged synchronously.
At present, the more mature explorations in the overseas market are still carried out around the traditional content structure.
Netflix brings narrative games to TV and makes mobile phones the controller; Roblox uses AI to lower the threshold for users to create games; Fable's Showrunner allows audiences to generate, rewrite and continue series; The entertainment universe built by Disney in cooperation with Epic is also based on mature engines, existing IP and continuous operation.
Netflix narrative game "Love Is Blind"
They are all increasing the interactivity of content, but are in no hurry to hand over all control to the AI large model.
The reason is actually very simple: what the audience really needs is not infinitely changing images, but characters, conflicts and stories worth entering.
If every choice can generate a new scene, it will be more difficult for creators to control the rhythm; if the character can answer any question, the character personality may quickly collapse; if the world has no boundaries, users often do not know what they should do.
Traditional film and television rely on editing to decide what the audience sees, and games rely on rules to tell players what they can do, while generative content has not yet found a corresponding order.
In fact, Netflix has not invested much in interactive film and television content for nearly ten years after "Black Mirror: Bandersnatch". After all, if users want stronger interaction in front of the TV or on their mobile phones, they can just play games directly.
Poster of "Black Mirror: Bandersnatch"
In contrast, some so-called "AI interactive film and games" in China are more like packaging multiple popular concepts for sale, but in essence they are just a more refined Galgame, and even the distribution channels are mainly Steam. AI-generated art assets, free-talking NPCs, branching plots, real-time video editing and world model generated environments originally belong to different levels of capabilities. Their simultaneous appearance in one product does not mean that the boundary between interactive films and games has disappeared.
A chatty NPC does not equal having memory and motivation; a real-time modifiable image does not equal that users can explore freely. Naming these capabilities "next-generation film and games" in advance is somewhat similar to the metaverse in those days — the product is not yet built, but the sales office has already opened.
Speaking of the metaverse, ByteDance's rumored world model plan this time has also made PICO interesting again. After the previous business contraction, ByteDance has not completely stopped XR R&D, and new systems, self-developed chips and new-generation devices are still under development.
PICO CLI 0.5.0
If the world model is finally integrated into PICO, it may indeed alleviate the long-standing most intractable content supply problem in the VR industry.
From this perspective, Zhang Yiming may believe in the "metaverse" more than Zuckerberg does today.
Zuckerberg is busy shifting resources to AI glasses, trying to make virtual information cover the real world; Zhang Yiming seems to be preparing to use AI to revitalize head-mounted displays and directly generate a world that users can enter. One hopes the metaverse will step out, the other still wants to invite users in.
Of course, PICO is only one of the potential outlets, and it should not take away the leading position of the world model itself. Even if the real-time generation speed is fast enough to run on head-mounted displays, dizziness, latency, spatial consistency, content security and computing costs will still appear at the same time. VR will not automatically solve the problems of interactive film and games, but will amplify every mistake made by the model in front of users.
PICO 4
In the next few years, the more realistic product will most likely still be a hybrid form: screenwriters and designers are responsible for the world view, characters and key plots, traditional engines maintain rules, spaces and states, and AI is responsible for generating dialogues, performances, transition scenes and local branches.
It may be more free than today's games, and more unpredictable than traditional interactive films, but it is still far from a world that can maintain long-term memory, run stably and continuously tell stories.
The world model is enabling images to have the ability to respond to users in real time for the first time, which is of course a real technological progress. What interactive film and games need to solve is never just to make the world move, but also to figure out why this world is worth entering.
This article is from the WeChat public account "Yiyu Guancha" (ID: yiyuguancha), author: HAL, published with authorization from 36Kr.