HomeArticle

Dialogue with Min Wei, Founder of Shadow Intelligence: Behind AMD's $8.2 Billion Acquisition of Li Feifei's World Labs, the inflection point of the 4D world model has emerged.

晓曦2026-10-07 10:31
Data is never a matter of volume, but a matter of dimensionality.

On September 28, a major piece of news broke out: AMD announced on its official website that it will acquire World Labs, the world model company founded by Chinese-American AI scientist Li Fei-Fei, for approximately 8.2 billion U.S. dollars in an all-stock transaction. The embodied intelligence and world model communities on both sides of the Pacific Ocean have been completely stirred up. 

Industry insiders exclaimed Amazing. World Labs was established less than two years ago, with only two products: the 3D scene generation platform Marble, and the world model Atlas. The company has not disclosed its revenue, let alone a mature business model. The technical route is far from converging: the 3D spatial intelligence faction, the non-generative latent space faction, and the platform infrastructure faction are all making their own efforts, and no one has been proven to be the standard answer. 

Controversy emerged almost simultaneously: some said it was a strong alliance, while others called it a betrayal of independent exploration by the "Godmother of AI". But from the perspective of investors, this acquisition may be logical, since the acquirer was already one of its early shareholders and has a clearer grasp of the actual operating situation of World Labs. 

Yingshen Intelligence, which belongs to the same world model track as World Labs, sees the underlying logic more clearly: the 4D world model has reached the critical point of Scaling Up. Just like the previous development path of large language models, World Labs needs greater computing power support for large-scale deployment, while AMD intends to lay out its chip ecosystem based on the 4D world model architecture in advance. 

Min Wei, CEO of Yingshen Intelligence, judges that at the inflection point of large-scale development, coupled with the catch-up of competitors such as GPT-6, World Labs no longer has time to conduct research in the small-step fast-running way as a startup company in the past, but needs to shift to large-scale industrialization. 

In June 2024, Min Wei, former head of Alibaba's robotics team, founded Yingshen Intelligence in Hangzhou, in partnership with Liu Yebin, a long-term appointed professor at Tsinghua University who has been deeply engaged in dynamic 3D reconstruction for more than 20 years. On the first day of its establishment, the team made a "non-mainstream" judgment: large language models process one-dimensional language, video models process two-dimensional images, while robots face a 4D world of three-dimensional space plus one-dimensional time, and intelligence that can understand the high-dimensional world cannot be trained with low-dimensional data. 

In the following two years, on the path of pursuing large-scale data, what Yingshen did was a process of "subtraction": the high-precision 4D acquisition device was reduced from a "cage" with 80 cameras worth millions of yuan to 4 to 6 ordinary cameras; two 8-GPU servers were reduced to one consumer-grade GPU; the one-day training waiting time was reduced to quasi-real-time generation of tens of milliseconds. Based on this, at the beginning of this year, Yingshen Intelligence moved the 4D world model into shoe manufacturing factories, enabling robots that truly understand physical laws to complete gluing and sole pressing work like humans; in September this year, Yingshen Intelligence's 4D live broadcast and 4D games made their debut at the China International Fair for Trade in Services; in the same period, Li Fei-Fei released Atlas, which follows the same route of generating 4D from sparse perspectives. 

The two routes converged independently on both sides of the Pacific Ocean, but their fates are different. The story of Atlas stays in the demonstration stage, while Yingshen has already deployed robots in shoe factories: with millimeter-level precision, 99.9% yield rate, only 5 team leaders are needed to manage 40 robots on a 45-person production line; in China, Yingshen Intelligence has obtained orders of hundreds of millions of yuan, and abroad, the first batch of hundreds of robots have been exported to Vietnam with orders of tens of millions of yuan. Adopting the business model of "selling hardware + selling Token", customers from Egypt, Turkey and South America have come to cooperate one after another. 

What kind of vote of confidence did AMD's acquisition of World Labs cast on this route? After the "Godmother of AI" joined a large technology group, do independent startups still have opportunities? How far is the 4D world model from large-scale implementation? After this acquisition, we had a conversation with Min Wei, and the following is the sorted out record of the conversation. 

8.2 Billion Dollars, What Exactly Is AMD Buying

36Kr: When seeing the news of this acquisition, many people in the embodied intelligence circle on WeChat Moments used the word Amazing to describe it. But to be honest, we have always had a question: the technology of the world model is far from converging, why did AMD place such a heavy bet at such an early stage? Why is it willing to spend such a large sum of money?

Min Wei: I think there are two reasons why AMD placed such a heavy bet. First, AMD was the major shareholder and main investor of World Labs in the previous rounds, and it knows much more about the progress of World Labs than we do. Second, in early September this year, World Labs released its new model Atlas. Atlas has a key breakthrough: compared with the previous generation of Marble, it can generate 4D data through cameras with sparse perspectives. Li Fei-Fei demonstrated a clothes-folding scenario: when you shake off the clothes, the picture can be traversed arbitrarily like "bullet time" and generate 4D in real time. We have already achieved this technology in April this year. 

At that time, we judged that after this technology came out, the problem of large-scale 4D data acquisition could be solved. In the past, 4D data acquisition required a lot of cameras and a long processing time, which could not achieve real-time performance, nor could a large number of cameras be deployed in real scenarios. When we achieved 6 perspectives, we keenly found that this technology can be massively deployed and collected in the industrial field and daily life. Once the threshold of massive acquisition is crossed, there will be a major scale-up immediately, the world model will be closer to implementation, and it will shift from the previous basic research stage to the stage of large-scale data and large-parameter models. 

So based on our own understanding of this technical route and previous work, I judge that after the release of Atlas, it also crossed the threshold of massive acquisition on the 4D route, just like us. AMD must have seen this too, so it quickly joined hands and acquired World Labs with a huge sum of money. It can be predicted that next, they will use an industrial and large-scale approach to collect massive videos in the real world with sparse perspective devices like Atlas, such as 6 perspectives or even 3 perspectives, then convert the sparse perspective videos into 4D, and use 4D data to train a more powerful 4D world model. 

36Kr: That is to say, it is after the release of the Atlas model in September that the application value of this route has truly achieved a major leap?

Min Wei: Yes, it fundamentally solves the problem of how to collect massive data. This is the first reason. The second reason is the release of GPT-6. After the release of GPT-6, there will be a problem: traditional visual models that use video as training data will face great squeeze. 

36Kr: Even so, isn't AMD's move too radical? It can completely continue to make strategic investments, instead of integrating the entire company into its own system at such an early stage. What do you think of the increasingly close relationship between the two sides?

Min Wei: The core judgment is that once the data generation technology can achieve sparse perspectives, it reaches an inflection point and can be scaled up. Later, big players like GPT-6 have caught up, and there is no time for World Labs to continue conducting research in the small-step fast-running way of startups or unicorns, it needs to shift to large-scale industrialization. 

36Kr: So you judge that the track has passed the 0 to 1 inflection point and is about to enter the 1 to 10 stage.

Min Wei: Yes, the technology is already ready. If it develops at its own pace, the remaining work will definitely not be as fast and good as with the huge capital investment of large technology groups. They must also see that since Atlas has released the sparse perspective technology, other large manufacturers will definitely catch up. So if they want to continue to seize the first-mover advantage, they need to scale up in advance within the year before other large manufacturers catch up, and the time window is very important. 

Three Leaps in Two Years, Solving Industry Bottlenecks with Technology

36Kr: Speaking of this, you once put forward a well-known view in the industry: the problem of data is never about quantity, but about dimension. This is quite similar to Li Fei-Fei's point of view. Why did Yingshen start working on 4D data so early?

Min Wei: From the perspective of first principles, the world itself is 4D. If we want to build visual AI and make it carry more laws of the physical world, we must first align the space-time dimension with the real world. However, pre-training based on videos is very difficult to support the laws of physics. 

Take wind as an example. Wind cannot be modeled in the pixel space, it is transparent and has no RGB information. But in 4D space, it is a fluid in space with zero transparency, which can actually be modeled. 

Another very typical problem we encountered before is the change of background and illumination. Changes of background and illumination belong to changes of key features in the pixel space, and the model can only eliminate this correlation by adding massive data, matching the same action with various different backgrounds, so that the model can learn from the data that the background is irrelevant and not causal. But if we work on 4D, this problem is very easy to solve. In addition, occlusion, contact and model penetration, which are very difficult to learn in the 2D world, are very clear in the 4D world. 

So the underlying reason from the dimension perspective is: all real physical laws are defined in four-dimensional space-time. It is very difficult to learn physical laws with data obtained by dimensionality reduction projection such as 2D pixels. 

36Kr: Yingshen achieved an early breakthrough in the 4D data field. Can you break down for us what technical bottlenecks you encountered when you first worked on this field, and how we broke through step by step?

Min Wei: For 4D data, the biggest bottleneck when we started our business in 2024 was the acquisition cost, hardware threshold, and training duration. In the early stage, our technology used about 80 cameras for high-precision 4D acquisition, which required two 8-GPU servers, and a few minutes of data took a day to train. At that time, we used the idea of mathematical iterative optimization for 4D generation, which could produce a small amount of high-precision data but could not scale. First, the equipment cost is very high, 80 cameras cost more than one million yuan; second, 80 cameras cannot be deployed in large numbers in real production environments. We can move experimental desktops and simple scenes into the acquisition cage, but the real scene is far more complex than that, it is impossible to move production lines and machine tools into the cage, and the cost is ridiculously high. 

So over the past two years, we have been committed to three things: gradually reducing the number of cameras, and gradually reducing the demand for computing power. The third is to reduce the time required for acquisition, training and generation, using mathematical methods of iterative optimization and repeated solution to shorten the 4D data rendering time from one day to real-time renderable. 

Last year, we reduced the number of cameras from 80 to more than 10, the acquisition and training time from one day to several hours, the number of servers from two to one, and the cost was greatly reduced. In April this year, we achieved a new technological breakthrough: the number of cameras was reduced from more than 10 to 4 to 6, an ordinary consumer-grade GPU can be used as the server, the time is compressed to the millisecond level, and quasi-real-time performance can be achieved. At that time, we first made a 4D live broadcast demo, which was exhibited at the China International Fair for Trade in Services in Beijing in September. 

In mid-September, we suddenly found that Li Fei-Fei released Atlas, whose technical route is the same as ours, using several sparse perspective cameras to generate 4D. Of course, whether her version can achieve real-time performance remains to be investigated. 

From this perspective, as long as 4D generation can be done with sparse perspectives, it can be applied to various scenarios in the real physical world: homes, workstations, workshops, restaurants. Adding 6 cameras costs very little, has low requirements for the environment, and supports non-intrusive acquisition, the collected objects will not feel it at all. This makes large-scale industrial acquisition feasible. 

36Kr: You just mentioned that the number of cameras has been greatly reduced. How was this leap achieved?

Min Wei: The previous 80-camera mode is essentially a mathematical formula, without historical information and no prior knowledge. Now we train a large model ourselves, and use this model and generative method to achieve 4D reconstruction and generation. 

36Kr: The industry commonly uses the Real-to-Sim-to-Real route, it seems that Yingshen does not follow this route.

Min Wei: We skip the Sim stage and go directly Real-to-Real: one model, input 6-perspective videos, and the output is 4D. What we exhibited at the Fair for Trade in Services is the state that can see 4D in real time. 

36Kr: This seems to be a very good solution. The Sim-to-Real gap has not been solved in the industry. But what problems will we encounter if we want to skip the Sim stage?

Min Wei: The core Sim-to-Real gap comes from Sim itself. If we skip Sim, there will naturally be no Sim-to-Real gap. What is the essence of Sim? It is to make various assumptions and simplifications in an artificially defined environment, and the gap comes from these assumptions and simplifications. 

If we do 4D feature extraction in the latent space, then use 4D features to generate 4D data, and then use 4D data for supervision, no matter it is the embodied model of robots or other models, to make the features in the latent space as 4D as possible and align with the real physical world, its hallucination problem and various non-causal correlation modeling problems will be much less. 

The biggest problem of the Simulator is Sim, because Sim must make various simplifications. The advantage of large models is a bit like the "grammar rules" in the era of large language models. Before the era of large models, why was the generalization ability of dialogue systems not strong? Because there were grammar rules, which limited the generalization ability. Later, large language models adopted end-to-end implicit learning in the latent space instead of explicit rule definition, which greatly enhanced the generalization ability. Sim can be regarded as an explicit grammar rule: everything in that environment is clear, artificially defined and rule-based, the rule itself becomes the biggest bottleneck of generalization, and the source of the Sim-to-Real gap. The simplification is reasonable in this scenario, but unreasonable in another scenario. 

How to simplify? Letting the large model perform implicit learning through the latent space is a better path. Just like grammar, is it written by linguists one by one, or learned by the model in the latent space by itself? So far, large language models have proved that only by learning implicit grammar in the latent space by itself can we obtain the optimal generalization ability that reaches the human level. Thus we have an inference: if we remove Sim in the latent space, directly model the 4D world in the latent space, and use 4D data for supervision, the learned physical laws will be closer to the real physical world and have better generalization ability. 

36Kr: If we follow this route, will we encounter any new problems?

Min Wei: The new problem is that 4D data is needed for supervision, and the volume must be