More and more "ghost stories" about embodied data acquisition are emerging.
Embodied intelligence does not merely face a general shortage of data, but has entered a phase marked by the scarcity of valid data, which will further reconstruct the business models of data collection companies.
There are an increasing number of well-known "horror stories" related to embodied data collection circulating in the industry.
A case that has been widely spread in the industry recently goes like this: a company purchased data in bulk, and half of its team was assigned to do the same task: data cleaning. After nearly a month and a half of hard work, the data was finally ready to be fed into the model for formal training. However, subsequent statistics showed that only 10% of the data was valid (the so-called validity rate refers to the proportion of data that can truly enter training after cleaning and alignment, and brings positive improvement to model performance).
Even the model trained with this 10% valid data performs worse than the previous version in some scenarios.
The company paid for the full amount of data, but only one-tenth of it was actually usable.
This is the real situation encountered by the industry including leading embodied data collection companies in the first half of this year. The outside world tends to simply attribute this phenomenon to "data fraud", but we believe this is not a problem of one party deceiving another. It is actually a practical pitfall the entire industry has stepped into after overestimating the value of data scale and underestimating the systematic cost of data quality.
At a time when the data boom is heating up rapidly, the industry does not lack general data, but lacks effective training data that can upgrade model capabilities. This viewpoint should become a new consensus across the sector.
Merely emphasizing the acquisition of data volume while ignoring the systematic nature of data for model training means that even with a huge amount of data, not every model company can make good use of it.
The requirements of embodied model training for the number of computing power cards and training cycles are also different from the past: a leading model company told us that taking 100,000 hours of data as an example, under the premise of sufficient computing cards, a complete round of training may take more than two months, which exceeds the expectations of many companies that cannot afford such a long waiting period.
If we break the myth of "data volume supremacy" and return to the starting point of embodied data collection, we should rethink several questions:
- What exactly is the high-quality data that embodied models need to collect?
- Why is it difficult to scale high-quality data?
- Is the industry really in need of more companies that only sell data?
Collecting only video data cannot support the next-generation embodied models
There is an estimate circulated among technical leaders of multiple companies in the industry: to build a truly usable embodied model, about 1 million to 2 million hours of training data is required.
It needs to be emphasized that this is an industry judgment rather than a verified law, but it is enough to explain why the competition around "how to quickly obtain millions of hours of data" has heated up rapidly in the past year.
In order to obtain human operation experience on a large scale, the ontology-free field collection route represented by UMI has begun to attract attention.
The core idea of this type of solution is not complicated: data collectors are no longer required to stand next to the robot, instead, people can directly complete operations in the real environment, and record action and environmental information through portable devices.
UMI itself adopts a combination of handheld grippers and vision systems, and converts human demonstrations into robot-learnable action representations through relative trajectories and other methods. Its value lies precisely in reducing the dependence of data collection on real robot ontologies, allowing human operation experience to enter the training link at lower costs.
Problems also arise from this point.
When the industry takes "large-scale collection" as its primary goal, the limitations of ego-centric visual data are becoming increasingly prominent.
The underlying logic of embodied data collection is to describe human actions as densely and realistically as possible, and retain the entire process of human interaction with the physical world in its original form.
However, vision-first ego-centric data seems to record a large amount of information, but in essence it first records the relationships between pixels.
It can tell the model what it sees, but it may not fully tell the model why the action is performed in this way, where the contact occurs, how much force is applied, and why the object changes in such a manner.
Hand occlusion, spatial depth, contact state, force sense and tactile perception will all become natural blind spots of pure visual input.
This does not mean that visual data has no value. On the contrary, work represented by UMI has proved that vision-led human demonstration data can learn dynamic, bimanual, fine and long-sequence operations, and can achieve a certain degree of generalization in new environments and on new objects.
The real question is: when embodied models continue to develop towards more complex and finer physical interactions, can a single visual modality still cover all the information that the model truly needs?
The answer is becoming increasingly clear: at the very least, embodied intelligence is not short of more video data.
Current embodied data collection needs to go through a transformation from superficial pixel imitation to underlying physical law learning. The core of this transformation is: what exactly is the high-quality data used to feed embodied models?
In model training, valuable data does not simply record "what action took place", but ultimately allows the model to learn why the action is effective and under what conditions it still works.
Traditional UMI-style data collection solves the problem of "how to obtain human demonstrations at low cost", but as models continue to develop towards higher-level physical interactions, data itself also needs to add new dimensions.
High-value embodied data in the future needs to meet three conditions at the same time: high fidelity, multi-modality, and diversity. These three conditions are not simply in a parallel relationship, but form a continuously narrowing screening process.
There are not many companies that can provide high-quality data;
Even fewer companies can provide multi-modal data at the same time;
There are even far fewer companies that can meet both high quality and multi-modality requirements while maintaining sufficient diversity.
The first requirement is high fidelity.
The data expression form and accuracy indicators need to be as close as possible to the dimensions of real interactions between humans and the physical world.
For example, the real contact surface between humans and objects is the skin, not the bones inside the human body.
When the collection system only records the movement of hand bones, it can describe where the hand is, but may not accurately describe where the physical contact between the hand and the object occurs.
This is also why as the requirements of embodied models for fine operations increase, the spatial geometric accuracy of hand movements is becoming more and more important.
Some teams in the industry are trying to further reduce hand tracking errors to more accurately restore the contact relationship between humans and objects. The specific error level depends on sensors, tracking schemes and usage scenarios, and the situation of the entire industry cannot be simply summarized with a single number.
The second requirement is multi-modality.
The value of multi-modal data is to enable the model to move from merely seeing actions to understanding interactions.
The language modality can provide semantic and intent information for long-sequence tasks, helping the model understand the task logic behind continuous actions; modalities such as tactile and force sense can provide contact feedback beyond vision, bringing more information for dynamic generalization and fine operations.
This is also why the ontology-free data collection route represented by UMI is undergoing new changes:
It is gradually evolving from a single visual trajectory to the fusion of multi-modal information including vision, touch and language.
At the same time, hardware is also evolving from partial hand collection to a more complete human operation perspective such as "head + hand".
The last requirement is diversity, because data is extremely prone to "inertia".
If the diversity of collected data is insufficient, the model may perform well on 4 to 5 limited specific tasks, but will quickly fail once the environment, object or operation mode is changed.
Therefore, truly valuable data does not mean repeatedly collecting the same action, but allowing the model to see a sufficient number of tasks, environments, objects, operation modes and failure scenarios.
Making efforts on the three dimensions of high fidelity, multi-modality and diversity is the only way for embodied data to get rid of the dominance of single visual data and move to the next stage.
Only by clarifying what high-quality data the model truly lacks can we further discuss "why high-quality data is difficult to obtain".
Why is it difficult to achieve both high quality and large scale?
After the ontology-free data collection route represented by UMI became popular, a new imagination emerged.
Since there is no need for robot ontologies, can data be scaled up quickly like internet text as long as we expand collection activities?
Facts have proved that it is not that simple.
Many people have underestimated the difficulty of expanding the capacity of high-quality embodied data, which involves the entire chain including data collection equipment, transmission, storage, cleaning, labeling, governance, and finally the access of data to model training.
At the underlying level, data quality depends on the ease of use of data collection equipment.
A good collection device should interfere with the original work of the collector as little as possible. In the ideal state, the collector should not even feel that he is wearing a data collection device.
This is the premise to guarantee the quality of raw data.
Just imagine, if the collection device is too bulky or has obvious intrusiveness, the collector will start to change his actions to adapt to the device.
Such unnatural interaction data distorted by the device may eventually cause the model to learn wrong operation methods.
At the same time, on the premise of not relying on large-scale laboratory motion capture systems and LiDAR, how to enable portable devices to achieve both sufficiently high spatial accuracy and sufficiently long working hours has also become a problem that must be faced in large-scale collection.
These two indicators often restrict each other.
Higher accuracy means higher requirements for sensors, computing power and calibration;
More modalities mean higher complexity of equipment, synchronization, transmission and data governance;
More complex equipment will in turn more easily affect the natural movements of the collector.
This is why high-quality data cannot simply copy the scaling logic of internet data.
Some high-degree-of-freedom dexterous collection devices still face challenges in stability during actual long-term operation.
In interviews, industry practitioners described some extreme cases to us: the device overheats and malfunctions after working continuously for a period of time, and even needs to stop for rest and maintenance midway.
These problems are not obvious during small-scale collection in the laboratory.
However, when data collection expands from hundreds of hours to tens of thousands of hours, device stability is no longer a product issue, but will directly become a business issue.
The damage rate of equipment not only reduces collection efficiency, but also brings additional costs such as hardware maintenance, replacement and personnel waiting time.
The real difficulty lies further ahead.
When the data scale continues to expand, the collection chain itself will face "traffic jams".
If an embodied company generates a large amount of data every day, every link from collection to upload, to storage, extraction, quality inspection and cleaning needs to be expanded synchronously.
The larger the data volume, the greater the pressure on storage and bandwidth;
The more modalities there are, the higher the difficulty of synchronization and governance;
Once the processing capacity of one link is insufficient, the collection capacity of the previous links will be wasted.
Therefore, the real scaling of the data collection industry is not as simple as deploying more collection devices.
It requires the entire chain to expand capacity together.
Especially under the crowdsourced field collection mode represented by UMI, data may come from different regions around the world, different personnel and different types of scenarios.
UMI solution demonstrates multiple task demonstrations
After tens of thousands of hours of data enter the system, completing collection, quality inspection, algorithm processing, cleaning and labeling, and ensuring that the time, space and semantic relationships between different data can be correctly aligned, is a complex engineering task in itself.
This also explains why when people's focus is blindly concentrated on "whether there is a lack of data", they tend to be overly tolerant of the invalidity of data.
The next stage of competition in embodied data may shift from the extensive "coarse screening" mode to more refined data governance.
What truly determines the value of data is no longer how much data has been collected, but: how much of the collected data can truly enter the training process and bring actual improvement to model capabilities.
This also means that data collection companies must bear increasingly heavy upfront investment.
For example, Jianzhi Robotics told us that to fully run through the complete commercial chain of ontology-free data collection requires capital investment of hundreds of millions of yuan.
This is also why the seemingly lightweight "field collection" business is actually a heavy-investment industry.
Calculate the cost account to see why high-quality data companies are scarce
No matter whether you attach importance to data scale or emphasize data quality first, model companies are essentially calculating an economic account.
Poor data quality naturally corresponds to low data prices.
The industry can actually distinguish junk data, otherwise low-quality data would not have to rely on price wars to compete.
Then why are there scarce companies selling high-quality data? Part of the answer can be found in the cost structure of data collection companies.
According to our multiple investigations, the highest proportion of cost in the current data collection chain is still data governance, followed by hardware, and then costs such as venues, operations and manpower. Moreover, this structure is still changing.
In the past, data was mainly video, but as model capabilities improve, data requirements have begun to expand from pure vision to tactile, force feedback and more modalities.
This means that the complexity of data governance will not disappear naturally as the data scale expands.
On the contrary, after moving from pure vision to multi-modality, work such as cleaning, alignment, quality inspection, labeling and compliance verification will become more complex.
But this does not mean that governance costs will rise forever.
What will truly determine governance costs in the future is the algorithm and automation capabilities of data companies.
If governance still relies on a large amount of manual labor, the larger the data scale, the more difficult it is to reduce costs;
If models can participate in data screening, quality inspection, alignment and governance, governance costs will have room for continuous decline as algorithm efficiency and computing power utilization improve.
This means that the next round of competition among data companies will not be about who has more manpower, but about who can process more data with fewer people.
This is the core reason why it is difficult for high-quality data to be both good and cheap.
Looking at hardware costs in contrast: devices such as sensors, collection gloves and head-mounted displays may gradually follow the supply chain logic of consumer electronics in