HomeArticle

Generate a 3D dream home from a single image, where you can freely swap tables, chairs and all kinds of furnishings as you wish, and another world generation model has been launched.

智东西2026-09-01 16:43
Application scenarios cover embodied intelligence, games, film and television, XR and other related fields.

Reported by AI News on September 1, today, Shanghai-based 3D generative large model company Yingmu Technology released the world generation model Hyper3D WorldGen, driving 3D generation development from "single asset" to "complete scene".

After users upload a scene image, the model can automatically identify the main objects in the frame and generate corresponding 3D assets. It also supports manual frame selection for target generation, and reconstructs the spatial and physical relationships between objects through the CAST architecture independently developed by the Yingmu Hyper3D team, finally generating a 3D scene composed of multiple independent, editable, replaceable and interactive assets.

CAST, short for "Component-Aligned 3D Scene Reconstruction from Single RGB Image", is a scene-level generative research result released by the Yingmu Hyper3D team in 2025. Its core function is to complete the 3D structure of occluded objects from a single image, judge the relationships of contact, support, suspension and others between objects, and place multiple independent objects into a unified, physically reasonable 3D space. This research once won the Best Paper Award of SIGGRAPH 2025.

It is worth noting that the generation capability of WorldGen has been applied to real production scenarios including embodied intelligence simulation training, film and television production, game development and XR, promoting scene-level 3D generation to evolve from visual display to industrial workflow.

Before the model release, media including AI News interviewed Wu Di, Founder and CEO of Yingmu Technology, and Zhang Qixuan, Co-founder and CTO, and exchanged views on topics such as the technical route of WorldGen, application progress and the trend of the 3D generation industry.

01. Generate scenes from one single image, assets in the scene are interactive

At the media exchange meeting, Zhang Qixuan demonstrated the generation process of WorldGen: after uploading an image of an office washroom, the system can automatically identify main objects such as taps, plates and fruits, generate 3D previews of single objects in about 2-3 seconds, and restore the complete scene in about 2-3 minutes.

All these objects are independent assets that can be selected, moved and replaced. After enabling the SimReady function, the system can estimate physical attributes such as object mass, size and friction coefficient for the assets, and fruits can roll into the sink under the action of force.

In addition, users can select different generation accuracies according to the importance of objects, quickly determine the scene structure first, and then refine the key assets.

Wu Di gave an example that if a generative model is used to directly convert an image of a meeting room into a 3D scene, the tables, chairs, walls and other objects in the frame are usually merged together, making it difficult to select and edit separately. In the generation process, WorldGen can split the main objects in the scene, generate independent assets respectively, and restore the spatial and physical relationships between them at the same time.

In order to balance scene availability and visual integrity, WorldGen adopts a hybrid expression method: objects that need to be edited and interacted with use mesh models with complete geometric structures to facilitate subsequent adjustment and addition of physical attributes; the background adopts 3D Gaussian Splatting which focuses on restoring the appearance of the scene to retain more environmental details.

02. Cover four types of application scenarios, take the lead in advancing in embodied intelligence and film and television fields

Zhang Qixuan introduced to AI News that WorldGen is not a model developed for a single industry, and its application scenarios cover fields including embodied intelligence, games, film and television production and XR. At present, WorldGen has promoted landing applications in different directions.

In the field of embodied intelligence, Yingmu Hyper3D has cooperated with Dige Robot and Moxianfei to launch a simulation training solution: WorldGen is responsible for generating 3D scenes that can be used for simulation, Moxianfei carries out simulation adaptation, and Dige Robot provides computing power and development tool chains.

Zhang Qixuan said: "Generative 3D can rapidly expand the types of objects and scene layouts, but it will not completely replace real data. Robot training is more likely to use a mixture of simulation and real data in the future."

In the film and television field, WorldGen can solve the problem that video models are difficult to maintain consistent space and object positions in multiple shots. Creators can first generate 3D scenes with clear spatial structures, adjust shots, move objects and modify layouts in them, and then hand them over to video models such as Seedance 2.5 to supplement character performance, materials, light and shadow and picture styles.

In the fields of games and professional 3D production, the independent mesh assets generated by WorldGen can be imported into digital content creation (DCC) tools and real-time engines such as Blender, Unity, PlayCanvas and Unreal Engine to continue material adjustment, animation binding, level editing and performance optimization. In July this year, Yingmu Hyper3D also announced cooperation with Unity China, trying to open up the process from 3D generation to real-time game applications.

In the XR field, WorldGen can restore reference images into editable 3D spaces, allowing users to observe the scene from different perspectives and adjust objects and spatial layouts, which can be used for interior design, space display and immersive experience.

Zhang Qixuan told AI News that the application of WorldGen in the above industries is currently mainly in the internal test or technical verification stage, and the team will determine the follow-up investment priorities according to the actual needs of customers.

03. Focus on rigid body generation, continue to evolve towards "building the world"

In the interview and exchange session, in response to the question about the technical boundary of WorldGen, Zhang Qixuan said that the physical attributes such as mass and friction coefficient generated by the system are not accurate measurements of real objects, which are mainly inferred for mass distribution, friction coefficient and collider by combining asset volume, visual information, relationships with surrounding objects, traditional estimation algorithms and visual language models.

Therefore, these results cannot be used as precision engineering data. At this stage, WorldGen mainly supports rigid body generation, and does not yet cover articulated assets, soft bodies and fluids.

Regarding the relationship between WorldGen and the "world model", Wu Di said that the world model includes both the action model oriented to decision-making and action, and the model that generates interactive and trainable environments, and WorldGen is closer to the latter. Yingmu Hyper3D will not define itself as a world model company for this reason, and its core positioning is still 3D generation.

Wu Di believes that video models are good at final picture presentation, while 3D can clearly save objects, spatial structures and physical relationships. The two routes do not replace each other, and are more likely to be combined in a hybrid way in the future.

From single asset to complete scene, WorldGen makes the generation results have stronger controllability, editability and workflow compatibility. As 3D generation evolves from "looking good" to "truly usable", scene-level generation is becoming a new basic capability connecting industries including games, film and television, embodied intelligence and XR.

04. Conclusion: 3D generation competition shifts to scene availability

The 3D generation capability of WorldGen is not limited to generating a single asset or a 3D space. It can construct multiple independent assets, understand spatial relationships and physical attributes, and form a complete scene that can be further edited, interacted with and operated. The focus of 3D generation competition is shifting from the visual quality of a single model to the controllability, physical rationality and workflow compatibility of scenes.

At this stage, WorldGen is still mainly focused on rigid body generation, and some physical attributes also rely on model inference. It is still in the verification and early landing stage in fields such as embodied intelligence, games, film and television and XR. Whether its subsequent value can be truly released depends on whether the generation results can stably enter the professional production process. However, evolving from "generating assets" to "building scenes" has become a highly feasible scheme in the 3D generation field.

This article is from WeChat official account "AI News" (ID: zhidxcom), written by Yang Jingli, edited by Li Shuiqing, and authorized for release by 36Kr.