HomeArticle

Zhang Yiming is personally building ByteDance's real-time world model

Tech商业2026-09-08 10:56
Zhang Yiming is leading the team in person, and ByteDance may launch its world model next month.

ByteDance founder Zhang Yiming is personally overseeing a real-time "world model" built on the company's Seedance video AI, which Bloomberg reports could launch as early as next month.

Zhang Yiming is no longer responsible for ByteDance's day-to-day operations. He stepped down as CEO in 2021, and later resigned as chairman, handing over the operation of TikTok, Douyin and other ByteDance businesses to others. However, according to a Bloomberg report on September 7, he is currently personally coordinating engineers from multiple business units to advance a project: a model capable of rendering an interactive three-dimensional world and allowing users to travel through it in real time, with all content sourced from the cloud.

The key lies in data. Bloomberg reports that the model has a latency of about 50 milliseconds and a frame rate of 20 frames per second. Such speed is fast enough to allow users' actions or voices to reshape the scene in real time without waiting for pre-rendered video clips.

Built on ByteDance's existing video generation system Seedance, the model is designed specifically for Pico, the VR headset that ByteDance confirmed it acquired in August 2021 without disclosing the purchase price. Pico is the target hardware for this model. In addition, the model is also oriented to the Douyin platform, where creators can produce live streams, short dramas and games on Douyin without purchasing expensive graphics hardware, as the heavy computing work will be completed on ByteDance's servers rather than on the headset itself.

Bloomberg interprets Zhang Yiming's personal involvement as ByteDance's move to secure a place in the artificial intelligence race, which already features two prominent figures: Li Feifei, a Stanford AI pioneer who co-founded World Labs to build spatial intelligence models; and Yann LeCun, former chief AI scientist at Meta, who has argued for years that language models alone cannot enable machines to truly understand the physical world.

A world model is not a 3D chatbot, but a trained system that can predict how a space and its objects will behave when objects move, similar to the predictive capabilities required for a robot to navigate a kitchen or an autonomous vehicle to identify a street.

This is the cutting-edge field ByteDance is chasing, and it is no trivial matter. ByteDance has set four top AI priorities for 2026. The world model ranks first, followed by consolidating Seedance's advantages in the video field, improving coding tools, and driving monetization of Doubao.

36Kr reported in June that ByteDance has invested three to four times more in world model training data (covering video, 3D and robotics data) than other domestic labs have invested this year. That is no small sum. It is reported that the Seed R&D team led by ByteDance's head of AI Wu Yonghui has set an internal goal of launching at least one world model by the end of 2026, and benchmarking it against Google's Genie 3.

Previously, ByteDance founder Zhang Yiming told his Seed AI team to avoid distilling competitors' models even if it sacrifices short-term progress. Earlier, the White House accused Moonshot AI of distilling Anthropic's Fable model to build Kimi K3.

Frankly speaking, this is an extremely challenging goal for a technology that no one has fully cracked beyond lab demonstrations, and all companies including Google and Meta have not fully mastered this technology.

The cloud bet may be more important than the model itself.

For anyone building AI infrastructure or consumer-grade hardware, this deserves special attention. Rendering an interactive 3D world usually means either strapping powerful GPUs to your face (as in the solutions of Meta Quest and Apple Vision Pro), or you can only get stuttering, low-fidelity images.

ByteDance is betting on bypassing this trade-off by rendering all content in the cloud and streaming it to local devices at 20 frames per second. If this solution works beyond demonstrations, it will shake the premise that immersive VR requires expensive local hardware to deliver a good experience. Pico's headsets are likely to be lower-priced, as will ByteDance's subsequent products.

This also fits ByteDance's existing profit model. Douyin's business is short videos and live streams, not game consoles or workstation chips. The cloud-rendered world model can be directly connected to the company's already large-scale profitable platform, without requiring millions of users to buy new hardware first.

None of this is set in stone. Bloomberg's report notes that the release timeline has not been finalized and could still be delayed. ByteDance has not publicly disclosed any information about pricing, availability outside China, and how Pico will integrate with content creators. But one thing is certain: the company's usually low-profile founder believes this is a project worth stepping back into the public spotlight for.

This article is published with authorization from 36Kr, originally from the WeChat Official Account "Tech Business", authored by Tech Business.