Who is profiting from the "tuition fees" in the robotics industry?
After doing housework for decades, Gao Bo, a full-time mother in her fifties, discovered for the first time that her movements could be monetized.
She lives in Shandong Province, and has to take care of her teenage son on weekdays, making it difficult for her to go out to work. Now, every day when Gao Bo cooks, does the laundry and cleans the house, she records her movements with her mobile phone. She told the media that she could earn 120 yuan for six hours of work every day.
These first-person perspective videos will be processed and used as training materials for robots to learn to do housework.
Gao Bo is no exception. On social media, similar scenarios are becoming increasingly common: in McDonald's, a girl wearing collection devices cleans the dining tables; in street food stalls, chefs record the angles of their wrists and movement trajectories while tossing woks to cook dishes...
Embodied data is stepping into the "crowdsourcing" era. In the past, robot data was mainly produced by collectors who remotely controlled real robots in training centers. Nowadays, with the popularization of non-entity collection devices such as first-person perspective cameras, data collection posts are no longer limited to training centers. Every place where people work can become a classroom for robots.
Behind this crowdsourcing boom is the huge data gap in the robotics industry. For robots to work in the physical world, they need to "supplement their brain" first, but the current situation is that data for understanding the physical world is severely insufficient.
A whole industrial chain has emerged rapidly around robot data collection and training, and the primary market has also poured in with great enthusiasm.
However, behind the bustle, the accounts are not easy to balance. Who will make money first in this industrial chain? How much of the continuously expanding data production capacity can truly be converted into robot capabilities and customer orders?
The entire industry is competing to "give lessons" to robots
According to data from Interact Analysis, by the end of April 2026, 64 data collection and training centers had been put into operation nationwide, and there were at least 90 projects including those under construction and in the planning stage, among which at least 13 centers deployed hundreds of robots.
In venues of thousands of square meters, home scenarios, supermarkets, warehouses and factories are restored on a 1:1 scale. Dozens or hundreds of robots line up to "take lessons", while collectors wearing VR headsets manipulate robotic arms to repeatedly carry boxes, sort goods and operate tools.
On recruitment websites, embodied intelligence data collectors have also started to appear in batches, and some positions do not require prior work experience.
A data collection service provider that previously served internet clients such as Baidu Maps added a team of more than 20 people at the beginning of this year, specializing in binocular collection of general data business. According to the relevant person in charge, two leading robot manufacturers and an elderly care real estate developer have taken the initiative to cooperate: "The demand is quite large, it can be said that we have caught up with the trend."
This boom comes from the shift of development priorities in the robotics industry. With the continuous advancement of robot bodies, joints and whole-body motion control, robots can already run, dance and box, and the industry has begun to devote more energy to models and practical tasks.
New bottlenecks have also emerged: robots can repeatedly complete one action in the training ground, but if they change an object, adjust the placement angle, or enter a room with different light and environment, the success rate may drop rapidly, which means the generalization ability of robots is weak.
Therefore, the robotics industry is trying to replicate the Scaling Law of large AI models: when the model architecture and training methods are relatively stable, the performance can be improved by increasing data, parameters and computing power.
This rule is still in the verification stage for robots, but it has already changed the industry's expectation for data scale.
According to industry estimates, the entire industry has accumulated about 500,000 hours of high-quality training data at present, while 100 million hours of training data may not be enough to achieve intelligent emergence. Calculated at this order of magnitude, the gap between the two is at least about 200 times.
After the demand increases, the industry still needs to solve the problem of where the data comes from. Traditional real robot collection is difficult to fill this gap. Some practitioners revealed that one trainer usually can only remotely control one robot, and one robot often requires two people to cooperate, and only dozens of real robot collections can be completed in one day, because robots "work" much slower than humans.
The non-entity collection that has heated up intensively in 2026 has lowered the threshold for expanding data production capacity.
This route emerged around 2024, and this year it has moved from papers and prototypes to complete sets of products. For example, the EGO device (note: first-person perspective data device) worn on the head records body movements, and the UMI (note: universal manipulation interface) held in the hand records the movement, rotation, opening and closing of the hand; wrist cameras and tactile devices supplement close-up pictures and contact information.
Technology opens up supply, and policies further amplify construction demand. In June 2026, the Ministry of Industry and Information Technology and the State-owned Assets Supervision and Administration Commission of the State Council launched a special action for real-scene practical training, requiring ten provinces and cities to select no less than 20 key scenarios respectively, and plan to form more than 100 high-value application scenarios by the end of the year, driving the landing capability of ten-thousand-unit scale.
With multiple forces overlapping, the data collection industry has quickly been crowded with companies from different backgrounds.
Hardware companies such as Orbbec and Tujian Technology are the most like "shovel sellers". Starting from 3D vision, Orbbec extends its capabilities in 3D cameras, calibration and large-scale manufacturing to EGO, UMI and wrist cameras; Tujian starts from flexible electronic skin, and uses tactile gloves to record contact force and force changes that videos cannot provide.
Entity and model companies such as Ubtech and Zibianliang have entered the data collection link for both equipment orders and training data. When local governments build data collection centers, they often purchase robots, teleoperation equipment and training systems in packages. The two projects won by Ubtech in Huizhou and Hohhot have a total amount of more than 130 million yuan. Zibianliang develops models, robot bodies and data collection tools at the same time, which can determine what data to collect in the next round according to the model performance, shortening the cycle of collection, training and testing.
Mifeng Technology, incubated by Agibot, tries to turn the data capability that originally served its own robots into an independent platform, undertaking collection, governance, training and evaluation externally, and acting as the organizer of the entire data production chain.
Jingdong and Synergy's advantages lie in their ready-made scenarios. Jingdong has logistics and retail scenarios as well as personnel organization capabilities, which can extend data collection to warehouses, shopping malls, factories and ordinary families; Synergy Intelligence already has a large number of robots running in warehouses and factories, hoping that these devices can generate data while working. The company disclosed that its robot brain has an installed capacity of more than 50,000 units, accumulating more than 500,000 hours of multi-form real robot data.
The whole industry is accelerating to "print textbooks" for robots, but more "textbooks" do not mean that students can learn well.
What kind of "textbooks" do robots need?
In June this year, robotics company XDOF shared an experiment.
The team prepared to teach the robot to fold T-shirts. After training with the ordinary imitation learning method, the robot succeeded 20 times in 20 tests. Later, the team continued to add more demonstrations that had completed the task, thinking that the increase of data would make the model perform better, but the success rate first dropped to 2 times, and finally all 20 tests failed.
The problem lies in those seemingly qualified videos: the clothes are indeed folded well in the end, but the process is mixed with pauses, hesitations, re-grasps and invalid adjustments. When humans watch it once, they think it is just a normal operation; but the robot will accept the whole process completely, and even take the hesitation as the standard answer to learn.
This is close to the perception of domestic practitioners. An industry insider told *NoNoise* that since this year, the demand for data collection has expanded rapidly, but a lot of data is found to be useless after being collected.
What the industry really lacks has gradually become the ability to judge what robots should learn from the data.
If a human movement wants to be converted into robot capability, the first step is to determine what to collect. Moving boxes pursues stability and rhythm, screwing screws requires precision and force control, and folding clothes also needs to deal with deformation, occlusion and multi-step operations. With different task definitions, the collection requirements for cameras, movement trajectories, tactile and joint states will also change accordingly.
These contents can all be converted into hours on the "capacity table", but they play completely different roles after entering the model.
After the collection is completed, the data will pass the first acceptance check: whether the task is completed, whether the picture is blocked, whether the sensor is offline, whether the image and control instructions can be synchronized, and whether human faces, mobile phone screens or trade secrets are captured.
A data service provider said that at present, they can ensure that at least one of every three or four collected data is valid, which is already relatively efficient in the industry.
After cleaning and labeling, the data has to face a more troublesome hurdle: when the same batch of data is in the hands of different companies, the acceptance results may be completely different.
Embodied data is deeply bound to the robot body and model architecture. The degrees of freedom of 6-axis robotic arms and 7-axis robotic arms are different, the motion spaces of two-finger grippers and five-finger dexterous hands are different, and the camera positions, sensor configurations, coordinate systems and control frequencies are also different. The same trajectory can be directly trained in one company, but when changing to another robot or model, it may need to be remapped, or even completely unusable.
Therefore, embodied data has at least three layers of validity standards.
The first layer is collection validity, which confirms that the task is completed and the files are available; the second layer is dataset validity, which confirms that the data has been cleaned, labeled and aligned, and can enter the customer's training pipeline; the third layer is model validity, which confirms that after adding this batch of data, the success rate, generalization ability and failure recovery of the robot are really improved.
At present, these three layers of "validity" are easily confused. When the supplier delivers 1000 hours of valid data, it may only complete the first two layers; what the customer expects is that the robot's capability will grow accordingly.
This also determines that embodied data is temporarily difficult to trade like ordinary commodities. Customers can hardly just place an order to buy 1000 hours of data. The more common demand is: a certain robot only has a 60% success rate in performing a specific task, and the supplier needs to judge which scenarios and movements are missing, and then design the collection, cleaning, training and supplementary collection scheme.
Therefore, validity can hardly become a fixed attribute of the data itself, it depends on whether the data can match a certain model, a certain robot body and a certain task.
Wang Xingxing, founder of Unitree, once pointed out that for robots, every input and output may produce deviation and loss. This is an important reason why the generalization ability and task success rate of current robot models are still insufficient.
This means that there is a data loss chain from shooting a movement to the robot really learning it.
As a result, a mismatch has emerged in the industrial chain: the previous costs have been paid, but the final training effect is still uncertain. There is no formula that robots will automatically become smart after being fed a certain number of hours of data.
When the data output cannot be directly converted into capability gain, the problem of the data collection business also arises: suppliers invest costs according to equipment, labor and collection duration, but customers are only willing to pay for the model effect. Who will bear the intermediate loss has become an account that the data collection industry must figure out.
The survival problem of the data collection industry
Looking along the industrial chain, the first money that is "received in the account" is the money for laying infrastructure. Cameras, gloves, grippers, teleoperation systems and robot bodies can be settled by unit or by set, and the training ground can also be accepted according to area, number of equipment and construction period.
Once entering the data trading link, the uncertainty increases sharply.
A number of embodied intelligence practitioners revealed to *NoNoise* that some local governments are willing to pay for robots and build training grounds, but they will require robot companies to repurchase data during project negotiations.
The *Embodied Intelligence Training Ground Research Report (2026)* jointly released by the Artificial Intelligence Research Institute of China Academy of Information and Communications Technology points out that training grounds are heavy asset investments. Although the sale of data products can generate revenue the fastest, relying solely on selling data is difficult to cover the heavy asset investment, and the payback period is relatively long.
Estimated based on public information, a professional teleoperator can only produce 2 to 3 hours of valid data on average after 8 hours of work. The quotation for domestic real robot data is about 500 yuan to 1000 yuan per hour, and some data collectors requiring technical background or on-site work have a monthly salary of 8,000 to 15,000 yuan.
Requiring data repurchase is equivalent to adding an "insurance policy" for the high investment cost.
But robot body manufacturers and model companies also have their own difficulties —
The cost of EGO and binocular devices has been declining all the way, and general tasks such as cleaning, storage and sorting are the easiest to expand production, and also the most prone to repetition. Different collectors wipe tables in different rooms, which does increase the total duration; if the changes of objects, movements and environments are limited, for the model, the newly added capabilities may not be as