HomeArticle

The "ChatGPT moment" for embodied intelligence may not arrive on the same day

吴怼怼2026-07-20 10:40
The data wall refers to the fact that high-quality real interactive data is expensive and scarce. The representation wall means that a sufficiently unified physical representation has not yet been formed among different tasks, scenarios and robot bodies. The closed-loop wall is manifested in the slow speed and high cost of trial and error in the real world, where a single failure may cause component damage and even affect the entire machine.

At the WAIC 2026 Embodied Intelligence Forum, the industry is exploring a system that can continuously learn and self-improve in the physical world.

After attending the event, a clear takeaway emerged: while "humanoid robots" remain the most camera-catching concept, the key terms repeatedly discussed on stage have shifted from degrees of freedom, joint counts, and motion demonstrations to data, models, simulation, deployment, failure feedback, and continuous learning.

This means embodied intelligence is entering a new phase.

Over the past two years, the industry first had to prove that robots "can perform tasks"; now, the question becomes whether they can perform those tasks hundreds or thousands of times consecutively, adapt to new environments, recover independently after failures, and continue improving their capabilities after deployment.

In other words, the competition in embodied intelligence is evolving from a race for better robot products to a contest over training systems and data infrastructure.

01. The Scaling Law for Physical AI Is Bottlenecked by Interactive Experience

The Scaling Law for large language models is relatively straightforward: more data, larger models, and greater computing power typically yield better performance.

But the physical world is not a web that can be directly scraped.

At the forum, Yao Maoqing, Partner of Intelligent Era and Chairman & CEO of Mifeng Technology, summarized the challenges facing physical AI as "three walls": the data wall, the representation wall, and the closed-loop wall.

The data wall refers to the fact that high-quality real-world interaction data is expensive and scarce; the representation wall means there is no sufficiently unified physical representation across different tasks, scenarios, and robot platforms; the closed-loop wall describes how trial-and-error in the real world is slow and costly, with a single failure potentially damaging components or even disabling the entire machine.

This assessment is highly significant.

When a language model makes a mistake, it usually just generates an incorrect answer; when a robot errs, it could shatter a cup, injure a person, or halt a production line. Therefore, what physical AI truly needs to scale is not just "data volume," but an executable, evaluable, and recyclable interaction closed loop.

This also explains why connecting a large language model to a robot does not equal achieving embodied intelligence. Large models excel at answering "what is this" and "what should be done," but robots must additionally solve "how to do it," "how much force to apply," "what to do after a failure," and "what changes occur in the world after the action is completed."

Moving from digital AI to physical AI means scaling embodied experience.

02. Data Is Shifting from a "Quantity Race" to a "Formula Competition"

The most notable shift at this forum is that almost no team still believes a single data source can solve all problems.

The currently emerging data structure roughly resembles a pyramid.

The base layer consists of the largest volume of internet videos and human behavior videos, which are low-cost and wide-coverage, helping models understand objects, scenes, and human actions. The middle layer includes first-person view, UMI, and other "platform-agnostic data," which records human motion trajectories via handheld or wearable devices to bypass the expensive data collection bottleneck of physical robots. The top layer comprises teleoperation data from real robots, actual deployment data, and post-failure recycled data — the smallest in volume but the closest to real-world execution.

During the panel discussion, Yao Maoqing estimated that if the "ChatGPT moment" for robots is defined as being ready to use out of the box, capable of understanding open-ended natural language instructions, and achieving a 70% to 80% basic success rate on common tasks, embodied data would need to reach the 100-million-hour level. This figure makes it clear that real-robot data collection cannot be the sole solution.

Su Hang, Associate Researcher at the Department of Computer Science of Tsinghua University, summarized the evolution of embodied pre-training into three directions: data is shifting from single-platform to multi-source mixing, models are evolving from behavior imitation to world modeling, and learning is transitioning from offline training to continuous evolution in real-world deployment.

This means future models will not only learn to "perform this action when seeing this scene." They will also understand task progress, predict action consequences, and distill experiences from different robotic arms, humanoid robots, and scenarios into transferable capabilities.

Data infrastructure is thus becoming an increasingly critical layer in the embodied intelligence industry chain. Mifeng Technology showcased the MEgo series of data collection terminals, the MEgo Engine governance platform, and the AGIBOT WORLD real-robot dataset, aiming to integrate collection, cleaning, annotation, evaluation, and deployment recycling into a complete workflow. Commercial value will likely shift from "selling a batch of data" to helping clients design custom data formulas and verifying that the data effectively improves model performance.

For embodied intelligence, scaling low-quality data can even exacerbate problems. A mislabeled action or a segment of misaligned sensor data will introduce execution errors during robot training. Data volume matters, but the ability of data to be reused across different robot platforms and verified in real-robot closed loops is far more important.

03. VLA and World Models Are Converging After a "Route Debate"

For some time, the embodied intelligence industry has seen a route debate: should robots directly use Vision-Language-Action (VLA) models to output end-to-end actions from images and instructions, or should they first build a world model to predict how each action will alter the environment before deciding how to move?

Based on insights from the forum, the two approaches are converging.

Intelligent Era's GO-2 highlights the chain of thought for actions, enabling robots to first make coarse-grained plans before completing high-frequency, fine-grained executions; Act2Goal achieves "think first, then act" by predicting intermediate visual states. Intelligent Era calls the integrated direction of VLA and world action models WRAM, short for World Reasoning Action Model.

Ren Zhiyi, Research Scientist at Physical Intelligence, presented the π0.7 model, which delivers another representative result: after learning to fold clothes on an ARX robotic arm, the model can transfer the skill to a UR5 dual-arm robot to perform similar tasks without fine-tuning for the new platform.

Cross-platform transferability is the key to building platform-level companies in the embodied intelligence industry.

If every new robotic arm, gripper, or robot requires re-collecting data and re-training the model, the robotics industry will remain stuck in the traditional automation project-based model. Only when models can map experiences from different platforms to a relatively unified action space can software and data unlock true economies of scale.

The direction proposed by Su Hang from Tsinghua University also points to unified representation: data evolves from single-platform to multi-source mixing, models shift from behavior imitation to world modeling, and training transitions from one-off offline learning to continuous evolution during deployment.

The ultimate physical AI system will likely be a multi-timescale system: the upper layer understands language and tasks, the middle layer predicts future states and formulates plans, and the bottom layer handles high-frequency control and real-time responses.

This is more akin to the coordinated work of the human cerebrum, cerebellum, and reflex system, rather than having a single model think about all problems at the same frequency.

04. The Real Watershed Is Shifting from "Video-Demonstrable" to "Robot-Usable"

The embodied intelligence industry once had countless stunning demos, but there is a huge gap between demos and usable products.

A robot successfully folding a piece of clothing once can make a widely shared video; what a factory actually cares about is whether the robot can work continuously, maintain a stable success rate, recover from anomalies, and have its maintenance costs offset by saved labor expenses.

At this forum, some newly released data began to carry industrial reference value.

According to on-site disclosure from Intelligent Era, its robots have been operating on a 3C production line in Nanchang, Jiangxi Province for six consecutive days, over 10 hours per day, completing nearly 65,000 operations in total with a 99.99% success rate.

Dyna Robotics revealed that its commercial foundation model DYNA-1 achieved a 99.4% success rate on the napkin-folding task, completing over 850 napkin folds in 24 hours without human intervention; for new tasks, it requires less than one hour of additional post-training data.

These metrics still need to be validated over longer periods, across more scenarios, and against real operational costs, but the shift in evaluation criteria is clear: the industry has begun measuring robots by continuous runtime, task success rate, required adaptation data volume, and human intervention frequency, rather than by the performance of a single demo.

On a deeper level, commercial deployment and general intelligence are not entirely contradictory.

In the past, people worried that working on specific projects would slow down general model R&D. However, teams like Dyna Robotics have proposed a "research-deployment flywheel," demonstrating that real-world deployment can expose unforeseen problems that cannot be observed in the lab and generate the most valuable long-tail data. Deployment is not just the end point after model training, but also the starting point for the next round of training.

The first to deploy robots at scale will gain access to real-world data earlier; the first to process that data efficiently will reduce subsequent deployment costs. This may become the most critical flywheel in the embodied intelligence era.

05. The "ChatGPT Moment" Will Not Arrive on the Same Day for All

At the end of the panel, guests offered varying timelines for the milestone: Yao Maoqing predicted two years, Zhao Zihao from Sunday Robotics estimated within three years, Ma Yecheng forecast four years, Ren Zhiyi suggested four to five years, and Zhang Zhenyou gave a range of three to five years.

Robots may not replicate the path of ChatGPT, which suddenly took the world by storm overnight. The "ChatGPT moment" that everyone refers to may not point to the same end point.

ChatGPT is a pure software product that can reach global users almost instantly once servers are ready. Robots, however, are constrained by manufacturing capacity, hardware costs, supply chains, safety standards, and on-site maintenance, making it impossible to achieve global distribution through a single webpage.

Physical AI is more likely to emerge in layers.

Imagine this progression: first, robots will appear in relatively structured scenarios with clear task value, such as factories, warehouses, and commercial kitchens; then expand to semi-open environments like supermarkets, hotels, and elderly care facilities; and the last scenario to mature will likely be homes, which have the most variables and the highest fault tolerance requirements.

Therefore, what truly matters in embodied intelligence is the ability to simultaneously build model capabilities, data infrastructure, simulation evaluation systems, robot hardware platforms, and real-world deployment networks.

In this competition, Chinese companies have advantages in dense manufacturing scenarios, complete supply chains, and fast engineering implementation speeds; the next question to answer is whether these scenarios can precipitate model capabilities that are reusable across different robot platforms and industries.

As more and more robots work continuously in the real world, and turn the successes, failures, and unexpected situations they encounter every day into new training materials, the intelligence emergence of physical AI is already happening — just at different speeds and in different places.

This article is from the WeChat public account "Wu Duidui," authorized for release by 36Kr.