Humanoid robots' ultimate competition comes down to their "brains"
From August 19 to 23, the 2026 World Robot Conference was held in Yizhuang, Beijing, with more than 300 enterprises bringing over 2,000 exhibits and over 150 new products making their debut.
Compared with a year ago, the conference has undergone tremendous changes: last year, robots on most booths were still performing show-like actions such as somersaults and boxing, while this year they are replaced by practical demonstrations including sorting medicines in a simulated pharmacy, picking up goods for customers in retail scenarios, and folding clothes and tidying up items in a simulated home setting.
However, on-site observation by reporters shows that robots still encounter problems from time to time in practical application scenarios.
On the afternoon of August 20, the reporter noticed that a long queue formed in front of an unmanned retail booth, drawing a large crowd of onlookers. A set of shelves were built on the booth, displaying both medicines and daily commodities, next to which was a checkout counter. Customers could place orders through a mini-program on their mobile phones. After receiving the order, the robot would pick up the goods from the shelf, place them on the checkout counter and complete the delivery.
One audience member placed an order for a bagged soymilk. After receiving the order, the robot first turned to the shelf to identify the position of the soymilk, extended its robotic arm to grab it, turned around and placed it on the checkout counter. The whole set of movements was coherent and smooth. As more and more people gathered to watch, another audience member placed an order for a small item placed on the inner side of the shelf. But this time, there was an obvious pause in the robot's movement. The staff explained that this was because the large flow of people nearby interfered with the robot's obstacle avoidance recognition. Some audience members commented that the robot got stuck just after picking up two items with such a small crowd. If it worked in a real retail store where there would definitely be more interferences than here, it would definitely not be able to function properly.
In fact, the above situation is related to the generalization capability of robots: at present, humanoid robots do not have a brain that can flexibly respond to various situations like humans. The behavior of robots relies on pre-stage data collection and scenario simulation training, which essentially executes according to pre-set programs. It can complete trained actions, but once an untrained situation occurs, it does not know how to deal with it.
"Customers buy robots to solve problems, not to serve them," Duan Yanbiao, CPO (Chief Product Officer) of Yunji Technology, told the reporter when talking about the challenges robots encounter in scenario implementation.
Yunji has been deeply engaged in the field of service robots for many years, and has deployed service robots in more than 40,000 hotels around the world. Duan Yanbiao said that the reason why customers have to "serve" robots in reverse is that robots will encounter a large number of long-tail problems during operation, which refer to unexpected situations that are low in individual occurrence probability but rich in variety. Taking hotel scenarios as an example, the long pile of carpets sometimes traps the wheels of robots, and similar situations are difficult to be covered in advance in laboratory and simulation training. Once it happens, customers need to spend time dealing with it specially.
Duan Yanbiao believes that these long-tail problems are one of the main factors hindering the large-scale commercialization of robots.
Wang Xingxing, founder of Unitree Robotics (688836.SH), said at the main forum of the conference on August 20 that the biggest bottleneck in the world at present is the insufficient generalization capability of embodied intelligence. After most current models undergo sufficient data collection and special training in fixed scenarios, their task success rate can be close to 100%, but as long as the operated objects are replaced or the environment is slightly changed, the success rate will drop significantly.
In Wang Xingxing's view, when a robot can complete about 80% of tasks in any unfamiliar environment, it is almost the ChatGPT moment for embodied intelligence, which means the emergence of a generational capability breakthrough similar to ChatGPT in the field of embodied intelligence. This time point will come in 2 to 3 years at the soonest, and 5 to 10 years at the slowest.
Against this backdrop, the competition centered on data accumulation for the robot "brain", model training and business model exploration has already kicked off.
01
"Spawning applications along the way"
At a highly deterministic fixed workstation, there is only a tiny gap between the operation success rate of robots and that of humans; once the task becomes longer and more complex, this gap will widen rapidly.
Xiaomi's new-generation humanoid robot "CyberOne" has been working at the nut workstation of Xiaomi Automobile Factory since March this year. The nut workstation is a fixed workstation on the automobile final assembly production line, where workers need to align nuts to screw holes one by one on the assembly line, tighten and fix them. Each round of operation lasts for a few seconds and is repeated hundreds of times a day with highly repetitive movements, but it has strict requirements for positioning accuracy.
Xiang Diyun, General Manager of Xiaomi Robot Division, told the reporter that the success rate of "CyberOne" when it first started working at the workstation was 90%; 4 months later, the success rate has increased to 98%; the success rate of humans completing the same operation is 99%, and this last 1 percentage point gap is expected to be caught up by the end of the year.
At highly deterministic fixed workstations such as the nut workstation, the efficiency of robots is highly correlated with task duration. Xiang Diyun said that at fixed workstations, the efficiency of robots can continuously approach that of humans, but once the task becomes longer, more complex, or there are variables in the environment that are not covered in training, the success rate will drop quickly. For example, for short tasks of 3 to 5 seconds, the efficiency of robots can reach 70% to 100% of that of humans; for medium tasks of 30 to 60 seconds, the efficiency of robots drops to 60% to 70% of that of humans; for long tasks of more than one minute, the efficiency will drop to about 30% of that of humans.
Wang Xingxing explained the technical reason for this efficiency attenuation in the above speech: every perception and control closed loop of the robot will produce deviations and losses, and the longer the task, the greater the accumulated deviation.
An investor who has long focused on the robot track told the reporter that in the robot industry, "99 points is equal to 0 points". As for what this missing 1 point means to users, we can get the answer by looking at sweeping robots. Sweeping robots have entered ordinary families for many years, with highly vertical tasks and relatively single functions. Even so, users need to do a round of preparation before starting the robot every time: pick up the scattered sundries on the ground first because the machine cannot recognize them; close the doors of certain rooms first because the machine cannot cross the threshold.
The investor believes that users buy robots to let the machines work for themselves. Once the machine cannot cope with slightly complex situations, users have to serve the machine in reverse. The more complex the task and the more open the scenario, the higher this cost will be.
Even sweeping robots are faced with such problems, so for general-purpose humanoid robots, the problems will only be more prominent.
At the site of this World Robot Conference, the reporter noticed that the home scenario is the scenario where various enterprises invest the most in demonstration, and various housework tasks such as folding clothes, cooking, sorting items and restocking refrigerators are demonstrated in turn. However, the home scenario demonstrations of all manufacturers are still in the prototype display stage, and no enterprise's products can be officially put into use in ordinary people's homes.
Wang He, founder of Galaxy Universal, said at the main forum of the conference that under the current technical conditions, directly allowing robots to enter ordinary homes is not a very good development path. Mo Lei, Vice President of Zhisquare, also believes that it will take at least 5 years for humanoid robots to enter home scenarios.
In addition, being able to complete a task and being able to complete a task well are two different things. At the conference site, the reporter observed that the robot of some enterprises takes about 10 minutes to fold a piece of clothing. In the robot table cleaning demonstration at a certain booth, a crumpled tissue stuck to the robotic claw, and the robot tried several times but failed to throw it into the bucket, finally the staff stepped forward to handle it. During the exhibition, there were even cases where multiple robots could not continue the demonstration due to on-site network interference.
Since the home scenario cannot be realized in the short term, the current industry consensus is to transition step by step from the B-end.
In this regard, the path given by Galaxy Universal is: first operate continuously in scenarios such as retail pharmacies and industrial production lines, then take institutional elderly care and hospital wards as intermediate transitions, then penetrate into rigid-demand home scenarios such as elderly care and disability assistance, and finally spread to ordinary families. Wang He calls this process "spawning applications along the way".
In the process of moving towards home scenarios, the capability of the robot brain needs to continue to evolve.
At the technical architecture level, the robot industry has experienced a direction convergence in the past year: in 2025, most embodied intelligence enterprises adopted the VLA model, namely the Vision-Language-Action model, which enables robots to process the seen images, received language instructions and output actions simultaneously; by 2026, most enterprises began to superimpose the world model on the basis of VLA, allowing robots to predict the future state of the environment before executing actions, so as to reduce the cumulative deviation in the process of task execution.
The industry has also formed an engineering consensus on the "fast-slow brain": the large brain is responsible for deep reasoning and task planning, while the small brain is responsible for high-frequency real-time execution and motion control. The division of labor between the two balances computing power consumption and response speed.
The architecture is converging, but to make the brain smarter fundamentally requires data. At present, the robot industry mainly relies on two paths to accumulate training data.
The first path is to deploy robots at customer sites to collect scenario data. For example, Zheng Xiaodan, head of JD Retail Embodied Intelligence Business, told the reporter that robots are moving from "demonstration" to "actual work". As a user with scenarios such as logistics, retail and healthcare, JD hopes to open up its own scenarios to jointly promote data collection with robot ontology manufacturers and model manufacturers. The goal in the next two years is to collect ten million hours of high-quality scenario data.
Zheng Xiaodan said that the measurement of data should not only focus on quantity, but more importantly, on high quality and scenario diversity.
The second path is simulation. The principle of simulation is to build scenarios in a virtual environment for robots to train repeatedly. The cost of the virtual environment is much lower than deploying robots at customer sites, and it can quickly create changes in various scenarios, such as changing light, moving the position of items, and simulating crowd interference. Theoretically, it can cover a large number of training scenarios at low cost.
However, the problem with simulation is that there is a gap between the virtual environment and daily scenarios. The physical laws in the virtual environment are simplified. For example, a tissue is a standard flexible object in simulation, but in daily scenarios, the touch of a crumpled tissue stuck to the robotic claw, and the sliding track of a wet and heavier tissue falling off the claw surface are very difficult for simulation to restore. Simulation can increase the total amount of data, but it is difficult to cover those occasional, random and unpremeditated variables in daily scenarios.
The industry calls this gap the sim-to-real gap, which refers to the transfer loss from simulation to reality. Robots trained purely by simulation often experience performance degradation after entering daily scenarios.
In Mo Lei's view, "actual work" is the scenario with the highest ceiling. Only by continuously working at customer sites can robots continuously expose problems, accumulate data and iterate models. He believes that an embodied intelligence company cannot only focus on models, but should integrate software, hardware and scenarios as a trinity. The brain and the ontology need to be iterated collaboratively. If a company only develops models without producing hardware or entering actual scenarios, the capability of the model cannot be continuously improved.
In terms of cost and hardware, Mo Lei told the reporter that if humanoid robots want to enter ordinary families, the hardware cost must at least drop to the price range of A-class vehicles (about 100,000 to 200,000 yuan). At the same time, a robot needs to work continuously for 8 to 16 hours a day and operate for 30 days a month without failure. The importance of hardware stability may be no less than the brain itself.
As a reference, the current service life of the dexterous hand of Tesla Optimus robot is only 6 weeks, the cost of a single hand exceeds 6,000 US dollars (about 42,000 yuan), and the annual cost of parts replacement for one robot alone is close to 100,000 US dollars (about 700,000 yuan).
At the exhibition site, the vice president of an embodied intelligence company used three sentences to summarize the current position of the robot industry to the reporter: data is the textbook, evaluation is the exam, and deployment is taking up the job. At this stage, most robot enterprises are still in the stage of learning with textbooks, cannot guarantee to pass the exam 100%, and are even further away from large-scale deployment in the future.
Chen Feng, relevant person in charge of ExoDynamics, said in an interview with reporters that 2026 is the first year of POC (Proof of Concept) landing for humanoid robots, not the first year of mass production, and large-scale mass production will not be realized until at least next year.
Chen Feng said that in the past, many robots had an operator holding a remote control behind them when performing on the booth; the smooth movements seen by the audience did not entirely come from the robot's own judgment. The most important step for humanoid robots to enter real scenarios is to "remove the remote control".
02
The "Brain" Business
Challenges such as insufficient generalization capability, uncovered long-tail problems, and unfeasible home scenarios all depend on whether robots can have a sufficiently smart brain.
A new business is taking shape around the robot brain.
Gao Jiyang, founder and CEO of Starinno, said at the main forum of the conference that the business model of embodied intelligence is changing, and it will be divided into three stages in the future: the current stage is dominated by complete machine sales, with hardware gross margin maintained between 40% and 60%; the second stage is the solution subscription stage, where customers no longer buy complete machines but pay for the overall solution, and the hardware gross margin will drop to around 20%; the third stage is the physical world Token sales stage, where customers pay according to the amount of intelligent capability consumed by the robot to complete tasks. The complete machine may even achieve negative gross margin, and the revenue source will shift to the continuous consumption of intelligent capability itself.
Gao Jiyang believes that this evolution path is similar to the process of the cloud computing industry shifting from selling servers to selling computing power. The brain will gradually evolve from an integral part of the complete machine to an independent source of revenue, while the hardware will in turn become the carrier that supports the brain.
For the brain to become an independent business, the premise is that the brain itself is smart enough. To make the brain smarter, two things are needed: data and computing power.
An analyst from a large securities firm in South China told the reporter that the current problem of the robot industry lies in the model, the problem of the model lies in data. If no large-scale company can emerge in the data link, it will be very difficult for the humanoid robot sector to see a major uptrend in stock prices.
Zheng Xiaodan said that data collection is a complete chain from front-end collection, cleaning and labeling to evaluation. The collection methods include multiple paths such as collecting through the robot ontology in scenarios, and collecting through first-person perspective devices worn by humans. JD is also expanding data collection in more industrial and home scenarios with external partners, but the standardization and large-scale development of the entire chain is still in the early stage.
In terms of computing power, Huifu Securities pointed out in a recent research report that Tesla completed the tape-out of its next-generation AI chip AI5 in April 2026. The computing power of a single chip reaches 2000 to 2500 TOPS (trillion operations per second), which is 4 to 8 times that of the previous generation AI4. AI5 is the computing power foundation of Optimus' next-generation brain, which directly determines how complex reasoning models the robot can run on the end side, namely on the robot ontology rather than the cloud.
Apart from data and computing power, Tesla has another advantage that most competitors do not have. Public information shows that Tesla FSD (Full Self-Driving System) has been included in the list of available regions in China in May, and is undergoing internal tests in cities such as Beijing and Shanghai. FSD and Optimus share the end-to-end neural network architecture and visual perception algorithm, and the autonomous driving data and algorithms accumulated by Tesla on roads can be directly used to train the robot's brain.
The above research report also mentioned that Optimus is expected to achieve a weekly output of 2,000 units by the end of the year, with the annual shipment target of 2026 reaching the 10,000-unit level, and the shipment in 2027 is expected to exceed 100,000 units. More Optimus