GPT-6 is here, we talked with multiple insiders about the moats and hidden reefs of embodied intelligence
Insiders in the embodied intelligence industry are beginning to worry that the future they have long awaited will be realized ahead of schedule, because the company that delivers this outcome may not be the one they assumed.
After the launch of GPT-6 Astra, this sentiment became particularly prominent in discussions at this year's Bund Summit. Zhu Zheng, co-founder and chief scientist of Extreme Vision, issued a reminder that language models have already taken over multimodal capabilities and are now encroaching on the embodied intelligence sector, holding full advantages in talent, computing power and capital. "The wolf is already at the door, abandon all illusions." In his view, the next one to two years will be a critical window, and the current embodied intelligence sector has not yet built real moats.
This concern is rooted in real-world developments. After OpenAI released Astra, a large number of demos showing it controlling robots emerged, and the embodied intelligence community quickly launched relevant tests and debates.
At the same time, Sam Altman has also clearly stated OpenAI's intention to develop humanoid robots and robots of other forms. For startups that are still polishing their first-generation products, the possibility that upstream model suppliers will become direct competitors has been placed on the table.
But at the same summit, we also heard another voice. Zhu Xing, CEO of Ant Group Lingbo, asked in a media interview: If OpenAI or Anthropic enters the autonomous driving field today, can they disrupt Tesla in a short period of time? Wang Qian, founder and CEO of Independent Variable Robotics, anchored the answer on data, verification and hardware: Do large model companies also need to supplement these infrastructures when they enter the physical world?
However, beyond the anxiety, some teams have already started to apply Astra to R&D: some use it to generate simulation scenarios, others try to make it control robots together with embodied models. The changes brought by large models entering the embodied field are far more diverse than the narrative of "who replaces whom". The absorption of general cognitive capabilities, the substitution of specialized models, and the takeover of a robotics company's business are three events with huge spans. Lumping them all into a single phrase of "dimensionality reduction strike" will easily mislead the public.
Frankly speaking, there does exist a clear boundary between embodied intelligence and general large models. But the boundary is shifting, and the problems that separate the two sides today will not automatically become permanent assets of any company.
01 Capabilities can be transferred, but physical experience still needs to be accumulated
Behind this round of anxiety lies a sharp inference: when general intelligence reaches a sufficiently high level, the capabilities required by robots will emerge naturally.
At the "Physical Intelligence: Divergence and Choices" roundtable of the Bund Summit main forum, Wang Qian from Independent Variable Robotics held reservations about this view. He pointed out that the improvement of language models' programming capabilities is closely related to the advancement of massive programming data, data infrastructure and verification systems. The progress in video generation has not directly solved the problem of robot motion control.
Cross-task transfer and capability reuse are of course important values of foundation models. Wang Qian's reminder is that every capability leap is often matched with corresponding learning materials, training targets and feedback mechanisms. If we remove these inputs from success stories and attribute all results to the natural spillover of general intelligence, we will misjudge the cost required for the next breakthrough. Historical experience cannot prove that capability leaps will never happen, but it is enough to remind the industry that the work behind such leaps cannot be omitted.
Therefore, after seeing the model control a robotic arm, the next question worth asking is: What kind of motion interface does it call, how much underlying control capability does it already have, whether it uses robot data, and what can it achieve after switching to a different set of hardware or a different object? Demos only show the end result of the whole system. To judge how much breakthrough the model has brought, we need to see clearly the composition of the entire system.
Xu Huazhe, founder of Shellbot Robotics, shared his team's tests on Astra at the same roundtable. He divided embodied tasks into three levels: semantics, space and physics. Under their test conditions at the time, Astra already had strong semantic and spatial capabilities: it could identify objects on the desktop, pick up thin chopsticks, and even stand the chopsticks upright and insert them into a cup.
Difficulties began to appear when it came to more complex contact tasks. Xu Huazhe mentioned that the team has not obtained ideal results for tasks such as picking things apart with chopsticks and folding clothes; only a few attempts at the two-handed object passing test succeeded. These observations cannot serve as a unified evaluation standard, but they provide a very clear dividing line: There is still a gap between knowing where an object is and where it should be sent, and mastering how things change after contact. Language and visual knowledge can provide prior support, but robots still need to handle variables that cannot be determined only by appearance in actual interactions.
A recent study based on 10 two-handed manipulation tasks from RoboDojo provides more specific observations of this relationship. In simulation tests, the GPT-6 Astra Direct control architecture achieved a success rate of 26% with an average score of 37.81; after combining π₀.₅ with Astra, the success rate reached 48% with an average score of 62.60.
But it is worth noting that in the hybrid architecture, Astra only corrected 14.4% of the execution control steps, and the remaining 85.6% directly adopted the motions from π₀.₅. The trajectory analysis disclosed in the study shows that Astra can correct target alignment errors, plan non-grasp operations such as pushing and toggling, and adjust subsequent actions when task conditions are not yet met; π₀.₅ provides direct motion reference and prior knowledge for object interaction.
This set of results demonstrates the value of combining general reasoning with local operation experience. Even a small number of intervention steps can have a huge impact on the success or failure of a task; the high proportion of underlying motions also proves that physical skills are still playing their role. The 10 simulation tasks are not enough to determine the victory or defeat of technical routes, and the 48% success rate also reminds us that there is still much room for improvement before a complete system can work reliably. This supports a direction of collaboration, and cannot be interpreted as large models have already solved all problems in embodied intelligence.
Back to the Bund Summit, Shen Yujun, chief scientist of Ant Group Lingbo Technology, proposed that robots must continuously receive sensor inputs, and even when reasoning or executing actions, they need to change their behaviors at any time based on new information. The physical world will not pause the falling of an object in its hand just because the model is still thinking. This puts forward requirements for the model in terms of continuous perception, timely response and deployment efficiency.
In fact, these requirements are also driving the technology choices of large model companies. When Google DeepMind released Gemini Robotics 2 in July this year, it provided embodied reasoning models, vision-language-action models and edge-side models respectively, and clearly pointed out that multi-fingered dexterous manipulation still faces challenges. Tech giants are entering the embodied field, and they are also configuring dedicated models and systems for embodied scenarios.
The capability boundary is being crossed, but the way to cross it is not necessarily that a single model does everything. Specialized training, system collaboration and edge-side deployment are all participating in this process. A more meaningful competition indicator is who can deliver these capabilities stably with less interaction experience and lower deployment costs.
02 Whether data can become a moat depends on who can continuously obtain effective feedback
In the final analysis, large models like GPT-6 are essentially the product of "overwhelming computing power applied to massive data", and their core competitiveness lies in the exploration and utilization of high-quality data. But can this capability be easily and directly transferred to the embodied intelligence field?
Zhu Xing used autonomous driving as an analogy: Why is OpenAI, which is so powerful, not entering the huge autonomous driving market to beat Tesla? We cannot ignore the gap between model capabilities and industrial capabilities. Leading training technology cannot eliminate the need for in-depth understanding of scenarios, nor can it directly bring a mature data pipeline. In his view, judging what counts as high-quality data and accumulating the relevant know-how is work with higher barriers.
This analogy reveals the real cost that cross-sector entry needs to pay. Of course, which market an enterprise chooses to enter also involves resource allocation, revenue expectation and strategic priority. The fact that it has not happened commercially cannot prove that it is technically impossible. Therefore, what we need to study is how many more links need to be supplemented to translate model advantages into industrial advantages.
For robots, the model is only one link in the delivery chain. What customers buy is a product that can continuously complete tasks. How the mechanical structure and sensors cooperate, whether the performance remains stable after equipment wear, who handles on-site anomalies, and whether maintenance costs can be controlled, all these will be included in the customer's cost accounting. Large model companies can build their own teams, make acquisitions, or cooperate with partners, but no matter which method they choose, they need to invest resources in all these work.
Wang Qian proposed at the roundtable that if OpenAI or Anthropic also need real-world data, simulation, evaluation and hardware infrastructure, then they are entering the same embodied intelligence competition.
But many people will question this judgment: If a company like OpenAI burns massive capital to accumulate data, who can compete with it? Indeed, this is a view worth pondering. When facing common challenges, perhaps not all parties will automatically have the same problem-solving speed. The computing power, talent and tool capabilities of tech giants may still significantly shorten the catch-up cycle.
Objectively speaking, the claim that "we own the data" needs to be subjected to more rigorous scrutiny. Data that can be repeatedly purchased on the market is not necessarily scarce; a large number of repeated successful trajectories in the same scenario cannot effectively cover failure cases either. Data scale has competitive value worthy of discussion only when it is linked to task distribution, feedback quality and model gain.
The visible reality is that the way data is generated is beginning to be affected by general models. Li Tianyu, co-founder and CEO of Source Policy Future, shared that his team has used GPT-6 to generate simulation scenarios through programming, and put humanoid robots into these scenarios to complete specified tasks. In their rapid actual tests, the generated results are complete, logically reasonable and practical. He believes that batch generation of 3D scene assets and the use of enhanced visual language understanding capabilities to assist data set cleaning can directly reduce the upfront data R&D costs.
This brings dual changes to data moats: links such as simulation asset construction and data cleaning may become cheaper, and some of the original production cost advantages will narrow accordingly. After scenarios can be generated quickly, the accuracy of the physical attributes of assets, whether they cover real failure cases, and whether they can improve the performance of real machines, instead require more judgment from the embodied team. There is still a gap between having production tools and knowing what to produce.
The accumulation that is harder to replicate may be a complete set of capabilities to continuously discover problems: knowing under which circumstances robots are prone to fail, being able to collect data in corresponding scenarios, identifying the truly missing information from failures, and then confirming improvements through training and evaluation. Deployment brings new problems, new problems drive model progress, and progress supports larger-scale deployment. A closed loop running in this way can make the first-mover advantage accumulate continuously.
03 The paradigm is still evolving, and the boundary will continue to shift
Frankly speaking, the widespread panic about GPT-6 in the industry today more or less stems from the technical paradigm inherited from multimodal models. The two are so much on the same track that the faster car behind will run over the one in front.
In Shen Yujun's words: If embodied models are still fine-tuned based on digital world models, for example, VLA based on VLM and WAM based on Video Gen, then once the foundation models in the digital world achieve capability leaps, these embodied models will be easily impacted.
Therefore, all current embodied intelligence manufacturers are clearly aware of the importance of so-called "physics-native" and "embodiment-native" models. On the one hand, this is to break the original technical dependence and find a unique solution for physical AI, and on the other hand, it is a kind of "track shifting" to some extent.
Recently, Embodied Intelligence Research Community interviewed some players taking different paths and paradigms, such as Shadow Intelligence and Embodied Brain Panshi. These enterprises that break away from path dependence are very calm when facing this issue.
When we threw the question "Can the brain-like route better resist the erosion of large models" to Zhu Senhua, CEO of Embodied Brain Panshi, he first corrected the wording: "I don't think it's about resisting."
In his view, neural networks inherently incorporate brain-inspired ideas. Emphasizing brain-like intelligence is to further explore theories and methods to overcome the bottlenecks of existing AI. He summarized these challenges as "data shortage, poor generalization, difficulty in incremental learning, and power consumption wall": how to reduce dependence on massive data, how to adapt to new scenarios, how to enable continuous learning, and how to control the power consumption of intelligent operation. "GPT has done nothing else but pile up data" "The fact that models become stronger does not mean that these four major bottlenecks have been solved along the way."
This answer brings the discussion to a deeper level. Will stronger AI in the future definitely follow today's model architecture? Continuous perception, complex contact, real-time motion and limited power consumption are all specific requirements put forward by the physical world for AI. To solve these problems, we may need to change the model architecture, training methods and even software and hardware organization. Embodied intelligence is also raising new research questions for general AI.
Another example is the 4D world model from Shadow Intelligence, which is also a type of exploration.
It tries to incorporate the three-dimensional space and its changes over time into the representation, to provide clearer spatial constraints for prediction and action. Min Wei, founder and CEO of Shadow Intelligence, reminded that embodied models obtained only by fine-tuning on the semantic base are extremely vulnerable to the dimensionality reduction impact of GPT‑6, "because it is a trivial task for large models".
At the same time, Min Wei also emphasized that future OpenAI may be more like an infrastructure provider similar to hydropower grids, and will not try to take over all application scenarios, which is determined by commercial boundaries.
Moreover, in a real, landed embodied intelligence product, the proportion of large model semantic capabilities is even less than half, and the remaining core value comes from complex systems such as physical law understanding, mechanical engineering and complex scenario interaction. "This is the underlying logic why Shadow Intelligence adheres to the native 4D model base, refuses to simply splice and transform on 2D video and semantic models, and prioritizes penetrating vertical scenarios such as light industry flexibility to get through the commercial closed loop," Min Wei said.
Back to the Bund Summit site, Shen Yujun gave a more open judgment on GPT-6 Astra at the roundtable: Embodied intelligence may even affect the development path