Observation on WRC2026: The "explicit line" is implementation, while the "implicit line" is data.
At this year's WRC, humanoid robots sort parcels on conveyor belts and tighten screws at workstations. Such "practical task" scenarios have become as common as the still bustling performance-oriented exhibitions, representing a gratifying progress of the industry for many people.
Looking back, however, this is not the whole story of WRC 2026, at least it is only the part of the "visible line".
"The amount of AI capability is determined by the amount of data, especially high-quality data." This sentence from Wang Xingxing, founder of Unitree Robotics, at WRC is almost the unspoken consensus of all exhibitors at this conference. The limbs of humanoid robots have been well developed, capable of running, jumping, flipping, and climbing up and down stairs, but their "brains" are still undernourished.
The question is, where does the data to feed the brain come from?
A set of data repeatedly cited at WRC shows that the current compliant data from real physical interaction scenarios in China is only 500,000 hours, while the commercial deployment of robots requires tens of millions of hours, leaving a gap of over 99%.
From 500,000 to tens of millions, there is a difference of two orders of magnitude.
As a result, "data" has become an invisible underlying line running through the entire WRC.
Players from all sectors show their unique strengths: some open source human behavior data, some sell data acquisition gloves, some bet on simulation, and others insist on real data from physical machines.
In this hustle and bustle, some embodied intelligence companies have no exhibition booths and are barely noticed, but they hold the most scarce batch of data in the entire industry.
Open Sourcing Human Data: Betting on the Mass Line
Guanglun Intelligence is one of the players that has taken the most aggressive actions on the data front at this year's WRC.
On August 20, they released EgoSuite-Open100K, the world's first 100,000-hour-level full-modal open source dataset of human behavior.
This dataset covers 100,000 hours of first-person human data, spanning 7 major categories of environments, 128 types of scenarios, more than 15,000 acquisition scenarios and over 15,000 tasks.
The data is mainly collected from the first-person perspective of a head-mounted device, with some additional wrist cameras, providing hand posture, full-body posture and event-level semantic annotation. The first batch of data is already available for download via Hugging Face.
Yang Haibo, CEO of Guanglun Intelligence, explained why they promote such a data route. He believes that relying solely on physical machine teleoperation can hardly support the supply of training materials at the level of tens of millions of hours, and the industry urgently needs to take two parallel paths: human video data and simulation synthetic data.
Therefore, Guanglun's strategy is the "mass line" — to open source human behavior data, build a Real2Sim2Real continuous learning closed loop, and even propose a 5-year 10-billion-hour embodied intelligence data co-construction plan. To put it plainly, since I cannot collect all the data alone, let the entire industry join in the collection work.
Parallel Development of Real Data and Simulation: Walking on Two Legs
At present, the combination of a small proportion of real data and a large proportion of simulation data is a makeshift solution amid data shortage, and has become the key direction that many industrial chain enterprises focus on.
Xinghaitu is a representative of this school. Gao Jiyang, its CEO, made a clear statement at WRC that they insist on giving priority to real data. He did the math: after combining three parts of costs — data cost, computing power cost and R&D engineer labor cost — "the data cost is actually relatively controllable, not as high as imagined. The most expensive part is the time of R&D engineers."
Xinghaitu's real data assets include the GOD real-scene dataset open sourced in September 2025, which ranks first in the world in download volume; it has expanded data sources by investing in companies such as Jianzhi and Yuanliu; it plans to expand the scale of real data to 1 million hours in 2026.
At the same time, however, Xinghaitu has not given up on simulation, and its paper citation volume ranks first in the world model field.
In a sense, Gao Jiyang's statement is more like an ideal. At present, simulation data is still the majority of embodied intelligence training data. In the future, everyone wants to use as much real data as possible, but the conditions are not yet mature now.
Even Unitree Robotics, the most popular star player, cannot avoid using massive internet data for pre-training, in addition to letting robots further adapt to the physical world through real robot operation data.
Walking on two legs — one responsible for enriching cognition of the world and broadening horizons, the other responsible for staying down-to-earth — will be the normal state of embodied intelligence development for a long time to come.
Obtaining Real Data Under Extreme Working Conditions in Mines
At this point, you may have noticed a problem: the data of all the players mentioned above, whether it is human behavior video, simulation synthesis, or physical machine teleoperation, almost all come from relatively standardized, structured, safe and controllable scenarios, such as laboratories, factory workshops, home environments and so on.
But there is a type of data that cannot be collected by any of the above methods.
That is the real physical interaction data under extreme working conditions — in mines 1000 meters underground, next to smelting furnaces with temperature of thousands of degrees Celsius, at the edge of blasting zones in open-pit mines, and in mixed operation sites with severe dust and light interference.
Human beings are unwilling to enter these scenarios (63% of mine accidents are caused by human factors), these scenarios cannot be accurately simulated (the physical complexity of unstructured extreme working conditions exceeds the modeling capability of simulation engines), and data cannot be collected via teleoperation (due to communication delay, occlusion attenuation, multi-vehicle channel congestion).
What Idriverplus holds is exactly this batch of data.
In the keynote speech at WRC, Hu Sibo, CEO of Idriverplus, unveiled a set of figures: as of the first half of 2026, the scale of Idriverplus's autonomous driving fleet exceeds 3400 units, of which the cumulative shipment of unmanned mining trucks exceeds 1900 units, covering nearly 40 mines around the world, and the maximum normalized operation scale at a single mine reaches 220 units.
What does this mean? It means that 1900 heavy-duty machines are operating in real conditions in more than 40 mines with extreme working conditions every day, continuously generating first-hand data of the full chain of perception, decision-making and execution. This is not clean data from the laboratory, not synthetic data from the simulator, not behavior videos shot by humans wearing head-mounted devices — this is data generated by machines working in real harsh environments.
Hu Sibo said a key sentence in his speech: "What autonomous driving has accumulated is not only algorithms and mileage, but also the cognition generated by 1900 vehicles running in more than 40 mines every day, and more importantly, the machine's understanding of the physical world."
The subtext of this sentence is that the moat of heavy-duty embodied intelligence does not lie in model parameters, but in the accumulation of real data under extreme working conditions.
Based on this batch of data, Idriverplus has built a three-layer heavy-duty embodied intelligence technology platform:
The bottom layer is a unified data engine that integrates multi-dimensional data of production, environment and vehicles;
The middle layer is a heavy-duty world model: the terminal action model is responsible for millisecond-level real-time decision-making, and the cloud training model keeps iterating;
The upper layer is a cluster decision-making and scheduling system, which can schedule more than 1000 agents in a single scenario.
More importantly, Idriverplus does not stop at the "transportation" link. 70% of mine operation links are non-transportation: drilling, blasting, excavation and loading. Idriverplus is extending the "brain" trained from transportation data to the whole process of operation: 34 mining robots have been delivered, the open-pit mine charging robot has realized full-process unmanned operation including hole searching, charging and vehicle moving, the smelting ladle handling robot transfers 75-ton slag ladles under high temperature of thousands of degrees Celsius, and the underground mining tunneling robot has completed the whole machine manufacturing.
The business fundamentals supporting all this are a set of figures that stand out prominently in the embodied intelligence track: the total revenue in the first half of 2026 reached 804 million yuan, a year-on-year increase of 97%; the gross profit was 212 million yuan, a year-on-year increase of 204%; the autonomous driving revenue increased by 107.3% year-on-year.
Against the background that the whole industry is generally burning money with high investment, Idriverplus is one of the few players that have already achieved a commercial closed loop and have a clear profit path.
Idriverplus has a clear positioning: it does not make terminal hardware, only develops the "brain", and acts as a technology partner for the intelligentization of heavy machinery.
This "brain supplier" model allows it to move nimbly, maximize the advantages of algorithms, data and cluster scheduling, while expanding along three lines: carrier, scenario and overseas expansion. It has cooperated with Proton Automotive and DeepMotion Technology to extend its business from mines to smelting and logistics parks, and bring Chinese solutions to resource-rich countries such as Australia, Brazil and Indonesia.
Data Acquisition Hardware: Selling Shovels for Mining Is Also a Profitable Business
If data is gold, then this school of players is selling shovels for mining.
At the WRC site, data acquisition hardware manufacturers broke out collectively.
Octopus Power displayed three products: fisheye headband, electromyography wristband, and exoskeleton isomorphic data acquisition glove, collectively called OctoSense. Its core selling point is the world's first realization of zero-sample generalization of electromyography across different individuals. The background is that the long-standing difficulty in traditional electromyography acquisition is that everyone's electromyographic signals vary greatly, and Octopus Power is trying to align human operation data with physical machines with almost no loss.
Zibianliang also set up an ontology-free data acquisition demonstration area at the site. Its three data acquisition hardware products unify the data output standard, and the collected materials can be seamlessly connected to its own data service pipeline.
Daxiao Robotics brought the 2.0 version of the ambient data acquisition solution, which uses ACE Ego Kit, ACE Data Engine and ACE Ego Matrix to build a complete link from data acquisition, automatic annotation to cross-ontology application.
BrainCo has a more comprehensive approach. Its dexterous operation data acquisition matrix integrates three data sources: physical machine execution, human demonstration and simulation generation — the dual-arm wheeled data acquisition platform collects vision, tactile and motion state synchronously, and the exoskeleton human data acquisition glove records the human operation process, converting human motion experience into data for robot learning.
Of course, there are some more cutting-edge practices.
For example, Fourier Intelligence proposed the "brain-computer data acquisition" mode, building a data system around the brain, human and machine. Its core purpose is to compare the EEG differences under the three states of "execution, imagination, and teleoperation of robots", so as to provide infrastructure for motor imagery research. BrainCo directly demonstrated at the site how the brain-computer interface connects human and robot: when a person generates a motion intention, the sensor collects the task-related EEG signals, converts them into executable instructions for the robot control system, and drives the humanoid robot to complete the corresponding actions.
In addition, Kankan Intelligence incorporated the popular smart glasses into its product line, and launched an ultra-lightweight heterogeneous data acquisition glasses of only 56 grams with vehicle-level accuracy, focusing on passive massive data collection in real scenarios.
No matter which method they adopt, the characteristic of this school is that they all try to build data acquisition infrastructure. The logic is very simple: you are all short of data, so I will sell you the tools to collect data.
Conclusion
At the WRC Developer Night, a five-tier data progressive framework was proposed: human demonstration provides knowledge prior, physical machine execution provides real-world experience, simulation improves exploration efficiency, failure data supports boundary breakthrough, and industrial scenarios provide a sustainable closed loop.
Among these five tiers, the scarcest and most valuable are the last two tiers — failure data and industrial scenario closed-loop data. Because the first three tiers can be scaled up through open source, hardware and simulation, the last two tiers can only be obtained through actual operation in real industrial sites.
From this perspective, the hidden data war at WRC 2026 has already shown different levels: players like Guanglun are competing for the breadth of data