HomeArticle

29 Embodied Intelligence CEOs Discuss "Generalization": Consensus and Divergence

明亮公司2026-08-27 17:29
The capital market no longer directly pays solely for the embodied "generalized imagination".

At the just-concluded 2026 World Robot Conference (WRC2026), "generalization" is almost an unavoidable issue for all enterprise executives.

According to incomplete statistics from *Suchbright*, around 30 enterprise founders or leaders on the main forum shared their views on generalization capabilities, which can basically be summarized as "two consensuses" and "one divergence": The first consensus is that generalization is the biggest technical bottleneck for current embodied intelligence; the second consensus is that the path to improve generalization still relies on data and scenarios; there are divergences on the technical route to achieve strong generalization.

Some guests believe that the current valuation of embodied intelligence is directly related to generalization capabilities, and stronger generalization capabilities also mean higher valuations.

With the listing of Unitree Robotics (688836.SH), the market no longer directly pays only for the "generalization imagination" of embodied intelligence. Research reports from a number of securities firms show that although generalization capabilities cannot be directly "quantified", generalization can be indirectly priced through four proxy variables: shipment volume, order quality, convergence degree of technical routes (stable valuation for converged links, high option value for non-converged links) and data flywheel (valuation premium flows to companies that can obtain real scenario data first).

Based on the speeches of guests at the WRC2026 main forum, *Suchbright* sorts out the views of enterprise leaders on "generalization capabilities", which can be roughly divided into six parts: definition, generalization bottleneck, technical route, data, implementation rhythm and divergence.

How to Understand Generalization?

Gao Jiyang, Founder and CEO of Starhaven Diagram gives the most "market-oriented" definition — generalization is training cost. He proposed that embodied intelligence moving towards productivity has three key indicators: speed, precision and generalization, among which generalization can be simply described as "training cost": hiring a robot requires training just like hiring an employee. High training cost means low generalization, and when the cost approaches infinity, the robot can never learn the task. At present, Starhaven Diagram can complete post-training for new and long-horizon tasks in about 10 hours, with the goal of reducing it to 1 hour or even completing post-training of long-horizon tasks with "a few pieces of data" next year.

Jia Kui, Founder and CEO of Transdimensional Intelligence splits general generalization into two coupled components: semantic generalization and physical generalization. Semantic generalization solves the problem of "the robot knows how to perform actions when seeing different objects", which is essentially human cognitive behavior intelligence, and the data must come from real human data; physical generalization solves the problem of "the robot can still work stably after the data during training and the physical conditions during testing change", which is solved by generative simulation and synthetic data post-training — this is also the first principle for Transdimensional to choose this route.

Hu Luhui, Founder and CEO of ZhiCheng AI proposes a three-dimensional framework: the generalization of physical intelligence includes three levels: task generalization, environment generalization and ontology generalization, and "these three are not additive but multiplicative", which leads to a surge in data demand and increasingly high requirements for quality and consistency.

Federico Pecora, Global Head of Physical AI Robotics Research at Arm from a system perspective, believes that advanced robots can achieve "cross-object generalization", especially in specific workflows such as sorting, logistics, warehousing and automobiles, which can cover most of the work; and promoting generalization research creates value at every step — it can be reused at the software, model and skill levels, and schedule resources across tasks, which is "a huge liberation for robot companies and reduces engineering transformation costs".

Why Generalization Is the Biggest Bottleneck

Wang Xingxing, Founder and Chairman of Unitree Robotics: The training collection success rate in fixed scenarios can be close to 100%, but "when the operated object or the environment changes slightly, the success rate drops very sharply". He attributes the root cause to the alignment deviation between the input and output of the AI model and the real world — the language model is digitally encoded with no loss in input and output, while every input and output of the robot has deviation and loss, "the last few centimeters or millimeters of error cannot be corrected". Therefore, he defines the "ChatGPT moment" of embodied intelligence as being able to achieve about 80% of tasks through voice or text instructions in 80% of unfamiliar environments, which will come in 2-3 years at the soonest and 5-10 years at the latest.

Cao Peng, Chairman of the JD Group Technical Committee and President of JD Cloud demonstrates from a large-scale perspective: in the standard tasks of JD Logistics, the efficiency and cost of robots are already close to that of human workers, but "it is impossible to replicate on a large scale to more and more complex scenarios because the model generalization is not good" — poor generalization directly blocks the road to large-scale replication.

Zhao Tongyang, Founder and CEO of Zhongqing Robotics lists generalization as one of the biggest bottlenecks: there are many customer demands, "if you have to train for each of them, there will be a problem of whether you are making a dedicated tool or a general one", the support of data, models and architectures for the simultaneous solution of multi-task and generalization capabilities is the fundamental bottleneck.

Xiong Rong, Chief Scientist of Zhejiang Humanoid Robot Innovation Center points out the common dilemma in the industry — "one task, one training": once the object, scenario and illumination change, data collection and training must be restarted; and the existing methods "do not really solve the problem of insufficient generalization", and even "lose the precise, reliable and efficient execution capabilities that the original robots have".

Han Zheng, Co-Founder and CEO of Sudo Technology puts forward the strictest business discipline: the classification of embodied intelligence should take "99% success rate" as the prerequisite, and those that fail to meet the standard are still in the scientific research stage; object operation should "meet 99% success rate and certain generalization at the same time", and "ensure the generalization of objects and environments first, and then pursue the generalization of object types", otherwise there is no commercialization path. He also questions the significance of "one task one training, piling up master and doctor students to collect data for iterative tuning" for generalization.

Generalization and Technical Routes

Chen Jianyu, Founder and CEO of Starbot Era, said: VLA "forms certain generalization but relatively limited" through imitation learning, and it is difficult to draw inferences from one instance without collecting data for new tasks; the construction of world models forms a common-sense understanding of the physical world (for example, water will flow down when pouring from a cup), which "has relatively good generalization and is expected to greatly improve generalization capabilities". He also distinguishes generalization levels: VLA can realize object generalization, but task-level generalization (zero-shot, new tasks that have never been seen or trained) is the "ultimate pursuit of generalization".

Zhang Yufeng, Founder and CEO of Unbounded Dynamics, believes: VLA pursues the probability distribution and muscle memory of "sufficiently large data volume and sufficiently good distribution", "emergence has never been observed"; the world model needs to break through the understanding of the underlying logic of the physical world to support scenario generalization. His implementation strategy is "practice skills in industry, practice generalization in commerce, and family is the ultimate challenge".

Wang Qian, Founder and CEO of Independent Variable Robotics, believes: relying on the strong generalization of the base model, it only takes 3 months to open up the whole chain of new scenarios (overseas peers took two and a half years), "better effect, lower cost and faster speed than dedicated models, expert engineering and single-point models is already a reality".

Xiong Youjun, General Manager of Beijing Humanoid Robot Innovation Center: the cerebellum side uses the cross-ontology generalized operation VLA model (the only one that has passed the national standard test), which has multi-configuration, multi-task and certain generalization capabilities, "new tasks can be unlocked with only 20 pieces of data".

Wang He, Founder and CTO of Galaxy General, although did not use the word "generalization", described the achievements: Galaxy StarBrain has realized "cross-ontology, cross-task and cross-scenario operation" — one model can be used as a handling robot for five-finger hands, and the action prediction loss can continue to decrease with human unlabeled data, and accordingly claims to reach the embodied intelligence ChatGPT moment in 2028.

Guo Yandong, Founder and CEO of Zhi Square pointed out: the human brain only consumes 20 watts of power to exert "very strong generalization capabilities", while the cloud computing energy efficiency of the most advanced AI chips is only on the order of 10⁻⁷ — the upper limit of generalization capabilities is strongly related to energy efficiency, which is the starting point of its brain-like architecture NeuroVLA (20-millisecond level response).

Zhu Zheng, Co-Founder and Chief Scientist of Extreme Horizon: hope the world model "can achieve generalization in general scenarios like a language model".

Song Bin, Co-Founder and General Manager of Beijing Fido Technology: to support generalization, the world model needs "multi-physics field coupled solution + long-sequence stable output + certain accuracy", "this one thing alone is estimated to take many years to conquer"; but he also agrees that the world model base itself should have strong generalization and general capabilities (analogous to nine-year compulsory education).

Generalization and Data

Generalization does not appear out of thin air, it depends on data. This part has the highest degree of consensus among enterprise leaders and is closest to the market pricing logic.

Cao Peng from JD Group pointed out that the industry recognizes that training a robot brain with "large-scale replication and good generalization" requires data of tens of millions of hours, while the current volume is only at the level of 100,000 hours; and there is the problem of "cross-ontology generalization and adaptation". JD's solution is to collect 10 million hours of human data + 1 million hours of real machine data in two years, and use the diversity of real scenarios to exchange for data generalization — as the data volume of its JoyAI-RA model rises from 1000 hours to 10000 hours, the generalization performance has improved significantly and exceeded π0.5.

Xie Chen, Founder and CEO of Wheel Smart, puts forward a grand framework: taking "10 billion hours of embodied data" as the goal to reach the basic generalization capability at the "high school graduation" level — real machines only account for about 0.1% of the data pyramid, and the rest rely on simulation and human data; its EgoSuite real-scenario human dataset is directly used to "improve the general generalization capability of robots".

Xu Jincheng, Founder and CEO of Pacini defines the premise of generalization from the perception side: generalization relies on "physical experience can be produced on a large scale → physical intelligence can be generalized on a large scale", and tactile sensation is the only stable modality not affected by the lighting environment (vision will fail with environmental changes), which is the perception basis for supporting generalization; its data "four-certificate" system ensures that the data is natively isomorphic and has high fidelity, thus supporting generalization.

Wang Xiaogang, Co-Founder of SenseTime and Chairman of Daxiao Robotics responded with the first-stage capability of the world model: data augmentation can "draw inferences from one instance to a hundred" — the world model generates diverse data as training supplement, to solve the generalization problem from the data side.

Jiao Jichao, Vice President of UBTECH and Dean of the Embodied Intelligence and Humanoid Robot Research Institute: cross-ontology transfer is a big challenge faced by the Transformer architecture — inconsistent motor parameters and configurations lead to "it works on this robot, but needs special adjustment on another".

Liu Wei, Vice President of Infineon Technologies, gives a solution from the brain-like side: function transfer and knowledge storage enable "no need to retrain for new tasks", reducing the demand for new data.

Du Dalong, Founder and CEO of Octopus Dynamics provides a generalization breakthrough at the data collection device level: its EgoBio mechatronic bracelet is the first in the world to realize cross-individual zero-data generalization — in the past, mechatronic/EEG devices needed to be calibrated and retrained for each individual, and about 60-70% of the audience wearing the device at the WRC site achieved generalization without calibration.

Generalization and Implementation Rhythm

Huang Qingqiu, Co-Founder and CTO of Moqi Intelligence puts forward the clearest "triangular contradiction": the three dimensions of capability are efficiency, accuracy and generalization — the industry pursues extreme efficiency and accuracy, while the family scenario pursues extreme generalization, the two are completely opposite, and logistics and service scenarios are in the middle. Therefore, "the family scenario is the end point rather than the starting point of capability accumulation": first cut into service scenarios such as hotel cleaning as a "family simulation training ground" (with sufficiently high generalization and sufficiently high data return), and then train a generalizable base model to gradually enter the family. The current generalization level of its model is "able to fold clothes with various patterns, but overcoats and suits are still not manageable".

Huang Yuanhao, Founder and Chairman of Orbbec has a consistent judgment on the rhythm: first implement in the dirty, tiring and heavy work that humans are most unwilling to do, "after the generalization is better and the intelligence is higher, then penetrate into more scenarios", and data, models and ontologies (hands, eyes, brains, sensors) need to be iterated simultaneously.

Hu Luhui from ZhiCheng AI supplements the opposite perspective: taking "rigid demands or tasks that people are unwilling to do" as the implementation scenarios, and believes that the family scenario is not more complex than the industrial scenario (the AI's understanding and causal relationship are similar), but the family scenario involves safety, privacy and ethics.

Xu Lei, Head of the Intelligent Robot Business Department of JD Group judges the convergence path from the data side: the existing data paradigm can already achieve "cross-ontology capability trained by data", and after the ontology converges in the future, "the model's general capability on the ontology and cross-scenario capability will be stronger".

Zeng Guang, CEO of Zhongke Cloud Valley gives the generalization practice on the industrial side: realize multi-task generalization with "integrated technology of perception, planning and control", and use the biped humanoid robot as a "generalized cross-terrain actuator" to access the industrial software system.

Yu Yinan, Founder and CEO of Weita Dynamics provides a framework for risk measurement after generalization: parametrically define the hazards of different scenarios with "failure cost" (empty plastic cup falling to the ground vs. glass cup filled with juice falling to the ground), products must undergo sandbox testing first, and be gradually authorized as capabilities improve.

Doubts and Divergences

Qian Dongqi, Chairman of Ecovacs Robotics put forward the most sharp question of the whole venue: the Scaling Law of language models is built on thousands of years of human language accumulation and Transformer token vectorization, but "no one has proved that the Scaling Law can be generalized to physical AI"; all robots in the exhibition hall "are designed for specific scenarios and cannot be generalized", which is a global challenge. Therefore, he puts forward a flexible solution — "carbon-silicon integration": users define robots, and human intelligence compensates for the gaps of robots, and gradually evolves from "tool → housekeeper → companion" before physical AI is truly established.

Du Dalong from Octopus Dynamics: adhere to "extreme scaling brings extreme generality", announce the roadmap of the world model (20 billion parameters in October, more than 100 billion parameters in the first half of next year), and propose that the evaluation itself must also be scalable, otherwise the model capability cannot be scaled.

This article is from WeChat Official Account "Suchbright" (ID: suchbright), Author: MD, published with authorization from 3