HomeArticle

Messi invested in Li Feifei, who then turned around and bought a robot "training ground".

AI前线2026-07-23 15:52
From 3D generation to physical simulation, Li Fei-Fei officially steps into the robot training segment.

Li Fei-Fei has appeared in Messi's investment portfolio.

One is the king of world football, the other is the "Godmother of AI" — two figures who previously had no overlapping paths are now connected through the booming embodied intelligence sector.

Following the conclusion of the FIFA World Cup in the United States, Canada and Mexico, football legend Lionel Messi has stepped into the "second half" of his career, scoring an impressive "goal" in the investment arena. According to the official Play Time website, this sports technology holding and investment platform founded by Messi in 2022 has added World Labs, a spatial intelligence company founded by Li Fei-Fei, to its investment portfolio.

Play Time is operated by a professional team, integrating VC investment and project incubation, with a focus on technology and media sectors. Messi does not directly participate in investment decisions, acting more as a core limited partner and brand endorser. Within its portfolio, World Labs is grouped alongside robotics firm FieldAI, physical-world foundation model company Perceptron, 3D tool provider Intangible, and AI data platform SuperAnnotate, highlighting Play Time's concentrated bet on the spatial intelligence track.

World Labs is responding to this expectation with concrete actions.

After completing a $1 billion financing round in February this year, the company began accelerating its upstream and downstream layout. Yesterday, World Labs announced the acquisition of robotics simulation company SceniX, which was co-founded by Yunzhu Li and other colleagues who worked together in the same laboratory during their Stanford years.

SceniX does not manufacture humanoid robots or compete in the full machine market. Instead, it builds virtual training and testing environments for robots, covering interactive data generation, policy training, performance evaluation, and fault and edge case testing before real-world deployment.

If the Marble model solves the problem of how to rapidly generate a world, what SceniX addresses is what happens next after a robot enters that world — when it pushes a cup or tugs on a rope.

The Football King Bets on World Models

In the past, the investors most closely associated with world models were primarily chip companies like NVIDIA and AMD, industrial software firms such as Autodesk, and specialized venture capital institutions.

Now, investment platforms centered around sports stars are also placing bets on spatial intelligence. This at least demonstrates that "world models" are evolving from an internal topic for a small number of researchers and tech firms into a broader capital narrative.

Certainly, the entry of celebrity capital does not prove that a technology has matured, nor can it replace products, customers, and revenue. What it truly brings first is attention, brand resources, and cross-circle communication capabilities. For a cutting-edge technology company like World Labs, this value cannot be underestimated, but it also should not be exaggerated as technical endorsement.

In addition, Play Time's current portfolio is also noteworthy, including:

  • FieldAI researches intelligent robotic systems that can operate in complex outdoor environments;
  • Perceptron targets physical-world foundation models;
  • Intangible provides 3D visualization and collaboration tools;
  • SuperAnnotate solves data problems for foundation models;
  • World Labs attempts to build world models capable of understanding, generating, and manipulating 3D spaces.

These projects do not form a strict industrial chain, but they all point to a major ongoing shift in AI: from understanding text and generating content, to comprehending space, predicting physical changes, and ultimately driving machines to act in the real world.

Therefore, the significance of Messi's investment in World Labs lies mainly in the fact that the spatial intelligence narrative, once confined to the tech community, is beginning to gain more diverse recognition from capital, while the challenges World Labs faces have also become more practical.

After completing the $1 billion financing in February this year, World Labs has obtained extremely abundant capital, while also shouldering equally massive commercialization expectations. The company has listed robotics, scientific discovery, creativity, and content production as potential directions for its world models. However, in public perception, its most representative product, Marble, remains a tool for generating 3D worlds.

Marble can generate structurally coherent, explorable, and editable 3D environments based on text, images, videos, and spatial sketches. This capability is valuable for gaming, film and television, architectural design, and virtual content creation, but it is difficult to single-handedly support World Labs' long-term industrial ambitions.

Robotics offers another possibility.

In robotics scenarios, the world model is no longer just a tool for producing digital content. Instead, it can become the infrastructure for training, testing, and evaluating intelligent systems. The former primarily generates watchable and editable worlds for humans, while the latter generates worlds for machines to repeatedly act, make mistakes, and verify results.

Acquiring a Physical Training Ground

As early as November 2025, World Labs had already demonstrated Marble's applications in robotic training and simulation.

Researchers used Marble to generate 3D environments such as kitchens, residences, and warehouses, then imported these scenes into MuJoCo, RoboSuite, and NVIDIA Isaac Sim for experiments on robotic arm manipulation, quadruped robot navigation, and industrial transport.

This proves that from the very beginning, Marble was not designed solely for content creation.

However, those early experiments also revealed its capability boundaries. Marble is mainly responsible for generating static spaces, visual backgrounds, and collision meshes. Objects that require actual robotic manipulation, such as cookware, boxes, conveyor belts, and microwave ovens, still need to be imported from other tools. Complex robot contacts, deformable object changes, and policy evaluations also rely on external simulation systems.

In other words, World Labs can already rapidly build robotic training grounds, but it has not yet fully mastered the physical laws within those grounds.

For ordinary 3D generation, sufficiently realistic visuals usually fulfill the main task. But for robots, visual realism is only the minimum threshold.

No matter how realistic a cup looks, if it moves out of nowhere before the robotic arm touches it, or if the rope passes through the gripper even though the robot is clearly holding it, this world has no training value.

Because what the robot learns is not just what an object looks like, but the causal relationship between actions and their consequences.

In June this year, Li Fei-Fei and the World Labs team published a long article specifically breaking down the systems currently collectively referred to as "world models" into three functional categories: renderers, simulators, and planners.

  • Renderers output the images that human eyes can see, with the core metric being visual realism;
  • Simulators output geometric, physical, and dynamic states, requiring the world structure and action consequences to be sufficiently credible;
  • Planners determine the next action an agent should take based on observations and goals.

Among the three, rendering is the most mature, planning is the most widely discussed, yet simulation is often underestimated.

Li Fei-Fei acknowledged in this article that Marble is only the company's first step into the simulation domain. It can simultaneously generate Gaussian point clouds for visual presentation and collision meshes for physics engine calculations, but there is still a long way to go before achieving a world model that can unify the processing of geometry, materials, dynamics, and motion planning.

SceniX is precisely filling this gap.

The company's public positioning is "hybrid simulation". It seeks to combine traditional physical models with learning-based models:

  • The physics engine provides interpretable geometric, mechanical, and causal constraints,
  • Neural networks learn the changes in deformable objects, complex friction, and dense contacts that are difficult to fully describe with fixed equations.

The team's background is highly concentrated around this technical path.

According to LinkedIn, the company has been established for only 1 year and 9 months. Apart from photos of its three co-founders, its official website offers little additional introduction.

SceniX co-founder Yunzhu Li earned his bachelor's degree at Peking University, then completed his doctorate at MIT under the supervision of Antonio Torralba and Russ Tedrake. During his postdoctoral research at the Stanford Vision and Learning Lab, he collaborated with Li Fei-Fei and Wu Jiajun. He currently serves as an Assistant Professor of Computer Science at Columbia University, focusing on robotics learning, structured world models, deformable objects, and multimodal perception.

Another co-founder, Changxi Zheng, is an Associate Professor of Computer Science at Columbia University with a doctorate from Cornell University. He has long studied computer graphics, scientific computing, and physical simulation, covering complex physical phenomena such as fluids, bubbles, thin rods, collisions, and acoustics.

Sonny Hu previously worked at Tencent and Amazon, leading projects including image generation.

In terms of division of labor, Yunzhu Li is responsible for robotics learning and world models, Changxi Zheng provides the foundation in computer graphics and physical simulation, and Sonny Hu and other team members drive engineering implementation.

Bridging the Reality-Simulation Gap

Specifically, SceniX targets the long-standing "Real-to-Sim" dilemma that plagues the robotics industry — the gap between the real world and simulation.

While current humanoid robot hardware is iterating rapidly, the evolution of intelligent capabilities is still constrained by physical testing. Real-machine debugging is costly and slow, and high-risk and extreme scenarios are difficult to replicate repeatedly. Simulation has thus become a necessary option for large-scale training.

However, the problem is that many simulation environments remain overly idealized — in simple terms, they are "too fake": the grasping, walking, and manipulation behaviors that robots complete in virtual environments may quickly fail in the real world due to minor discrepancies.

Yunzhu Li once expressed a similar view in an interview: "The hardware is already in place, and AI is advancing rapidly, but we are still training robots in overly simplified virtual worlds."

In 2025, the SceniX team collaborated with Columbia University and Google DeepMind to publish the paper *Real-to-Sim Robot Policy Evaluation with Gaussian Splatting Simulation of Soft-Body Interactions*.

The research reconstructs digital twins of deformable objects from real videos, restores the visual environment using 3D Gaussian Splatting, and tests tasks such as plush toy packing, rope threading, and object pushing. It verifies that performance variations of different policies in the simulator can accurately reflect their actual hardware performance.

This transforms the role of simulation, turning it into a filtering and validation system before real-machine deployment — first eliminating a large number of unstable model versions, then submitting a small number of candidates for testing on real robots.

After that, the team continued to advance along two technical paths.

  • The Interactive World Simulator uses real-machine interaction data to train world models that can respond to actions in real time, enabling the simulation environment not only to evaluate policies but also to generate new training data.

  • The latest PGRD (Learning Physics-Guided Residual Dynamics for Deformable Object Simulation) combines physical models with neural network residuals to characterize complex deformations such as ropes and plush objects at a lower data cost.
  • The former expands the data production capacity of simulation, while the latter improves its credibility for complex physical processes. SceniX's technical boundary has gradually extended from "replicating reality and validating policies" to "continuously generating robotic experience".

    Li Fei-Fei commented that SceniX's capabilities go far beyond lab demonstrations — its technology has been validated in actual deployments on real robotic hardware.

    The complementary relationship between the two sides has thus become clear.

    Marble excels at generating diverse, explorable, and persistent 3D spaces, expanding