HomeArticle

PrimalVerse (Yuanhao Dynamics) has closed a seed financing round of several hundred million yuan, and its foundational 4D world model is developed by the top-tier team from Tsinghua University.

光源资本2026-07-23 14:37
One of the very few teams in the world that have full mastery of the complete technical chain of 4D world models

Recently, 4D world model company PrimalVerse announced the completion of a seed financing round of several hundred million yuan, jointly participated by well-known financial investment institutions such as Lenovo Star, Gingko Valley Capital, Qifu Capital, and ZHUOYUAN ASIA, with strategic investment from Jifon Medical, a leading industrial player in the surgical robot field. This round of financing will be mainly used for the R&D of the foundational 4D world model and team expansion, and continuously promote its application verification in vertical scenarios such as embodied intelligence, gaming and film, autonomous driving, and medical surgery.

PrimalVerse was founded by a top world model team with Tsinghua University background. The team has long been deeply engaged in 3D/4D reconstruction, neural rendering, physical simulation, and robotics, and is one of the few teams in the world that can connect the complete technical chain of 4D world models. As a partner, Lighthouse Capital will continue to support PrimalVerse in advancing the technical R&D and multi-scenario application verification of 4D world models.

Defining the Fifth-Generation Foundational Model: From Generating Content to Deducing the World

Large language models have enabled AI to learn to predict the next word, and video models have enabled AI to start predicting the next frame. However, as AI delves into real physical world scenarios such as robotics and autonomous driving, models need to further answer:

What will happen to the world after an action?

Looking back at the development of foundational models, the first generation is represented by large language models such as GPT, which mainly solve language understanding, reasoning, and dialogue (text Token); the second generation integrates image Tokens and text Tokens, represented by Nano Banana and GPT Image 2.0, realizing image-text understanding and image generation; the third generation, represented by Sora and Seedance, introduces video Tokens and starts to handle temporal dynamics and action continuity. The first three generations of models mainly model language and two-dimensional visual content and have not truly entered the three-dimensional physical world. The fourth-generation native multimodal large models with 3D tokens will begin to represent geometric structures and physical properties; in the next stage, foundational models need to further integrate space, time, physical laws, and action interactions into a unified representation.

Based on this judgment, the PrimalVerse team proposes: The fifth-generation multimodal large model will be a world model foundational model with native 4D World Token as the core (4D = 3D space + time).

The 4D World Token can be understood as a "new language" for the model to understand the physical world: it not only records a single frame but also continuously represents what the objects in the scene are, where they are located, how they move, what physical properties they have, and how they will change under the effect of actions.

When language, images, videos, and 3D/4D world states can be jointly trained in the same model, AI will no longer learn only the visual appearance of the world, but the structure, laws, and action consequences of the world itself. The model will also move from "generating a video" to "running and deducing a world".

The founder of PrimalVerse stated:

"4D is not an additional capability of the world model, but the basic expression of the physical world itself. The first-principles problem of the world model is whether it can accurately characterize the transition of the world from one state to the next, which requires simultaneous understanding of space, time, physical properties, and interaction relationships. We believe that judging whether a model truly understands the physical world should ultimately align with 4D representation; only by first building a native 4D world model can we provide a unified and credible physical world benchmark for different technical paths."

PrimalVerse hopes to directly build the native 4D world state from the bottom layer, simultaneously learn space, time, physics, and interaction in a unified model, and incorporate geometry, material, motion, contact, and physical properties into the state change process, defining a 4D world model paradigm truly oriented to physical AI.

The World's Rare Full-Stack 4D World Model Team

Building a native 4D world model foundational model requires integrating four capability stacks: reconstruction, simulation, rendering, and robotics. The absence of any capability will limit the complete construction, continuous evolution, and real interaction of the world. Globally, teams that can cover and connect the complete technical chain are extremely rare.

The core team of PrimalVerse comes from top universities such as Tsinghua University, Peking University, Shanghai Jiao Tong University, and the Hong Kong University of Science and Technology. They have long been deeply engaged in key technologies of 4D world models and have formed internationally leading achievements in all core links of 4D world models:

In the direction of real-time neural rendering, SlimmeRF won the Best Paper at 3DV 2024; in the direction of controllable simulation, MARS won the Best Paper Runner-up at CICAI 2023, which is also the world's first open-source highly realistic autonomous driving simulator, creating a modular neural rendering autonomous driving simulation paradigm; in the robotics field, Dexora (ICRA 2026 Best Paper Finalist) is the world's first open-source VLA model for dual dexterous hands with high degrees of freedom; Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots (RSS 2026 Best Paper Finalist)

realizes the automatic generation of manufacturable, collision-free robot facial mechanical structures from a single portrait, reducing the professional design time from 22.8 hours to 11.7 minutes; in the direction of reconstruction and representation, TRELLIS.2 won the Best Student Paper at CVPR 2026, which is the world's first single-image-generated 8K 3D model, realizing the native 3D generation of geometry and materials for the first time, breaking through the industry's common reliance on 2D Diffusion for material generation.

These achievements together form a clear path:

Understand the world, construct the world, deduce the world, and ultimately drive agents to act in the world.

The PrimalVerse team has produced dozens of top conference achievements at CVPR, ICCV, SIGGRAPH, 3DV, RSS, ICRA, ECCV, etc., and has Best Paper-level achievements in the four underlying technology stacks of reconstruction, simulation, rendering, and robotics. The establishment of PrimalVerse is the first time the team has gathered years of accumulation in different frontier directions into a unified goal - to build the fifth-generation multimodal 4D world model foundational model.

Starting from Vertical World Models to Verify General Foundational Model Capabilities

The value of the fifth-generation 4D world model will first be released in scenarios with the highest requirements for dynamic space, physical realism, and action interaction.

In the embodied intelligence field, PrimalVerse has carried out cooperation with many leading enterprises in segmented fields, promoting joint verification around simulation data generation, action consequence prediction, strategy training, and Sim2Real. Among them, PrimalVerse and Geekplus jointly developed the world model Gravity 4D for real warehousing operation scenarios, which has been verified in end-to-end operations such as multi-SKU grasping, bin handling, and mobile robot collaboration. The team's accumulated capabilities in dynamic prediction, force feedback modeling, and high-degree-of-freedom manipulation in achievements such as Dexora and TA-VLA are also continuously transformed into world model solutions for real robot tasks.

In the medical surgery field, strategic investor Jifon Medical and PrimalVerse are carrying out in-depth joint R&D in directions such as scenario-specific world models for medical scenarios, embodied intelligent surgical robots, and intelligent simulation training platforms, deeply integrating world model capabilities with surgical robot bodies, clinical scenarios, remote surgery systems, and doctor training systems, promoting the evolution of surgical robots from precise execution tools to intelligent surgical infrastructures with scenario understanding, risk prediction, and collaborative operation capabilities.

In the gaming and film field, PrimalVerse has reached a deep strategic cooperation with a leading film and television company, and is simultaneously advancing the intention of cooperation on virtual studio and director projects. Based on achievements such as Light-X, the company introduces industrial-grade 3D asset generation, controllable camera movement, and controllable lighting capabilities into the actual production process, serving dynamic content production, virtual shooting, and editable scene construction.

In the autonomous driving field, PrimalVerse has a complete capability stack from highly realistic simulation, world state prediction to city-level 3D reconstruction: MARS supports controllable simulation of complex road scenarios (the first open-source highly realistic and controllable autonomous driving simulator, laying the foundation for autonomous driving neural rendering simulation, adopted by many large manufacturers), OmniNWM realizes multimodal, long-time-series world prediction and strategy evaluation; TideGS, jointly developed by the team and Great Wall Motors, breaks through the video memory bottleneck of large-scale 3DGS training, supporting more than 1 billion Gaussian Primitives on a single 24GB GPU, which is about an order of magnitude higher than the industry's single-card training scale (compared with World Labs Spark 2.0, which is about 100 million), providing underlying support for city-level native 3D world construction.

Embodied intelligence, medical surgery, gaming and film, and autonomous driving are not four completely independent business lines, but verifications of a set of 4D world model capabilities in different scenarios: autonomous driving provides a complex dynamic world, embodied intelligence provides action and physical feedback, gaming and film verify high-quality world generation and real-time rendering, and medical surgery tests the model's understanding and deduction capabilities in scenarios with high safety requirements, complex deformations, and fine interactions. Multi-scenario achievements and industrial cooperation will also continuously precipitate data, models, and system capabilities to feed back the general 4D foundational model.

In the future, PrimalVerse will start from the training, simulation, and content production requirements in vertical scenarios, gradually build a general 4D world model that can be migrated across scenarios and industries, and become the foundational model entry for physical AI to understand, train, and deduce the world.