HomeArticle

Valued at 1.2 billion USD in just three months: Amid the robot gold rush, the first to strike it rich are those who "sell motion data"

新芒X2026-09-09 12:18
For the robotics industry, the real opportunities may not lie only in manufacturing robot bodies.

The humanoid robotics startups that most easily grab public attention are usually those whose robots can run, dance, fold clothes, and even perform martial arts moves.

Behind these glamorous demo videos, however, a far less glamorous business is heating up rapidly: hiring people to pick up cups, fold clothes, and arrange tableware over and over again, then convert these movements into data that robots can learn from.

In September 2026, XDOF, a robot training data company, became the latest capital focus on this track.

According to TechCrunch, XDOF is in late-stage negotiations for a new round of financing, which may be led by 8VC, with the company valued at approximately 1.2 billion US dollars. It has been less than three months since the company exited stealth mode in June this year.

It needs to be emphasized that as of the release of relevant reports, this round of financing has not been fully completed, and the financing amount and specific terms are still subject to change. Therefore, the "1.2 billion US dollar valuation" is more accurately described as the price being discussed in the capital market at present, rather than a finalized result.

But investors' eagerness is not groundless.

Reports say XDOF's annualized revenue has approached 50 million US dollars, serving about 20 clients including some undisclosed cutting-edge AI labs. Back in June, the company just announced the completion of 70 million US dollars in financing, with investors including Thrive Capital, Spark Capital, Andreessen Horowitz, Lux Capital and WndrCo.

Why can a company founded only in 2024, which does not produce humanoid robots nor release foundation models, be pushed to the threshold of unicorn status in such a short period of time?

The answer may be: when everyone is competing for the "OpenAI" of the robotics era, XDOF chooses to become the "data factory" of the robotics era first.

I

What robots lack is not knowledge,

but "physical experience"

The rise of large language models is largely built on the massive texts, images and codes that the internet has accumulated.

Models do not need to write a book in person or actually run a store to learn languages, knowledge and reasoning patterns from the digital traces left by humans.

The world that robots face is completely different.

There are countless texts and videos on the internet about "how to fold a shirt", but robots need to know far more than that.

They need to judge the material and shape of the clothes, calculate the motion trajectory of the robotic arm, control the gripper force, identify whether the corner of the clothes slips, and re-grasp when the action fails. What the camera sees, how the joints move, when the gripper closes, and whether the action is finally successful all need to be recorded synchronously.

Ordinary videos can only tell robots "what humans did", but may not directly tell them "how the machine should move".

Therefore, a valuable piece of robot training data usually needs to include images, depth information, joint states, robotic arm trajectories, gripper opening and closing states, action instructions and task results at the same time. Some fine-grained tasks also require information such as tactile sensation, force and spatial position.

What is more tricky is that these data are difficult to crawl directly like web text.

To obtain a reliable robot operation trajectory, it often requires a real person to control the robot to complete the task; if the action fails, it is also necessary to rearrange the items, calibrate the equipment and perform the operation again. Different robots may have different sensors, mechanical structures and control interfaces, so the same set of data may not be directly reused.

This turns robot training data into an expensive physical production activity.

Language models consume data assets that already exist on the internet, while robots require humans to produce a new set of "motion internet" about the physical world for them.

XDOF is targeting exactly this gap.

II

What XDOF sells is not action videos,

but a "data factory"

XDOF was founded by researchers from the University of California, Berkeley including Philipp Wu and Fred Shentu.

Before founding the company, the two co-founders participated in the development of GELLO, a low-cost robot teleoperation system. It uses 3D-printed parts and general-purpose motors to make control devices with structures similar to the target robotic arm.

When humans push the small controller, the real robotic arm can complete the corresponding actions synchronously. According to the GELLO paper, the material cost of a set of control equipment can be kept under 300 US dollars, and it is more suitable for collecting complex, contact-intensive operation data than low-cost solutions such as VR handles.

This research later became the technical starting point of XDOF's business model.

The company's data business can be roughly divided into three levels.

The first level is to perform teleoperation directly on the client's robot.

XDOF deploys equipment and operators to collect data in the robots, scenarios and tasks specified by clients. This type of data matches the client's real products the most and has the highest commercial value, but it has high costs and slow expansion speed.

The second level is to use standardized robots and teleoperation equipment to mass-produce relatively general-purpose operation data.

These data do not necessarily serve only one client, and can cover common skills such as folding, grasping, placing, sorting and wiping, which are closer to the general training materials required by robot foundation models.

The third level is to collect human first-person action data.

Collectors wear cameras or motion sensors to complete daily tasks such as packaging, folding clothes, and organizing items. Although these data are not generated by the robot body, they can expand the diversity of tasks and scenarios at a lower cost.

What XDOF really sells is not just the finally generated data files.

Robot data collection involves hardware maintenance, operator training, task design, equipment calibration, data synchronization, quality inspection, failure annotation and format conversion. Errors in any link may make hundreds of hours of data unusable.

Therefore, what clients purchase is a complete set of production capabilities ranging from robots and operators to data cleaning, annotation, evaluation and delivery.

XDOF defines itself as a "robot infrastructure company" rather than a pure data annotation company. It has also launched the Foundry data platform for clients to browse samples, manage uploaded content and call full datasets.

In June this year, XDOF cooperated with academic institutions to open the ABC-130K dataset. According to its public data page, the dataset contains more than 130,000 dual-arm robot operation trajectories with a total duration of about 3590 hours, covering information such as images, robotic arm joint states, gripper states and task instructions.

Opening this batch of data ostensibly helps the robotics research community, but in effect acts as a product sample: letting the outside world see what XDOF can produce first, then selling larger-scale and more targeted data services to commercial clients.

III

50 million US dollars in annualized revenue,

showing who is already willing to pay

What attracts capital the most to XDOF is not just the story of "robot data", but the reportedly nearly 50 million US dollars in annualized revenue.

This figure means that before humanoid robots generate large-scale commercial revenue, data service providers in the upper reaches of the industry have already found clients willing to pay.

Why don't robotics companies and cutting-edge AI labs complete these tasks on their own?

First of all, data collection is a work that extremely consumes operational capabilities.

Enterprises not only need to buy robots, but also need to build sites, maintain equipment, train operators, and ensure that data generated at different times and by different equipment remains consistent. Labs may collect hundreds of hours of data, but scaling up to thousands or even tens of thousands of hours becomes a different kind of organizational capability.

Secondly, the robotics industry is still rapidly trial-and-error.

A startup company may test multiple robotic arms, sensors and models at the same time. If the data team has to be rebuilt every time the technical route is changed, the fixed investment will be very high. Outsourcing to XDOF is equivalent to converting part of the fixed cost into procurement cost that can be adjusted with the project.

Thirdly, the quality requirements for robot data are far higher than that of general image annotation.

XDOF once found in internal experiments that after adding more "successful demonstrations" to the clothes-folding model, the model's performance decreased instead. The reason is that some ostensibly successful data contain invalid movements, repeated grasps and pauses in the middle, and the model treats these noises as correct strategies.

This means that more robot data is not always better. Which actions are effective, where the failure occurs, and whether the task is truly completed all need to be judged in combination with the robot and the model.

What clients purchase is not cheap labor, but the understanding of the data production process.

However, the figure of 50 million US dollars also needs to be treated with caution.

Relevant reports use "annualized revenue" rather than audited annual revenue, and do not specify how much of it is stable subscription revenue, how much comes from one-off data projects, hardware delivery or customized services.

If a company's revenue mainly comes from building exclusive collection systems for a small number of large clients, the current revenue can grow very fast, but it may not have the high gross profit and replicability of software companies.

Roughly calculated based on the 1.2 billion US dollar valuation and 50 million US dollars in annualized revenue, XDOF's valuation is about 24 times the current revenue run rate.

Clearly, capital is not valuing it at the price of an ordinary outsourcing company, but betting that XDOF can eventually transform labor-intensive data collection into a standardized platform, resellable data assets and a continuously operating data network.

IV

"Scale AI of the robotics era",

may also only be a phased business

XDOF is often described as "Scale AI for the physical AI era".

This metaphor is easy to understand: Scale AI helped tech companies collect, clean and annotate data in the wave of large models, and XDOF hopes to occupy a similar position in the humanoid robotics era.

But the robot data business may be more complex than large model data services.

The first challenge is that clients may eventually choose to produce core data on their own.

For general-purpose AI labs and leading robotics companies, real operation data is not only training material, but also the most important competitive asset. If robots continuously generate data in factories and homes, companies may prefer to keep collection, model training and product deployment in the same closed loop, rather than handing it over to third parties for a long time.

XDOF's window of opportunity may appear in the early stage of the industry: most companies are in urgent need of data, but do not have the ability to quickly build a complete data production line.

But as leading enterprises expand, there is still no answer as to whether these clients will continue to outsource or bring data capabilities back in-house.

The second challenge is that the degree of data reuse between different robots is limited.

Language models can learn the same text together, but robotic arms have different lengths, degrees of freedom, gripper structures and sensor configurations. Data collected for one type of robot may need to be re-adapted when migrated to another type of robot.

If cross-embodiment learning cannot mature for a long time, XDOF will need to continuously produce customized data for different clients, and its business model will be closer to project services rather than "selling one copy of data repeatedly".

The third challenge is that data collection still relies heavily on a large amount of hardware and labor.

Building warehouses, buying robots, repairing equipment, and training operators all consume capital. The faster the business grows, the more sites and personnel may increase simultaneously.

If revenue growth must be accompanied by cost growth at a similar speed, it will be difficult to obtain the scale effect common to platform software. Whether XDOF can increase its gross profit margin depends crucially on whether it can reduce labor input through automatic quality inspection, unified data formats and standardized equipment.

The fourth challenge comes from alternative technologies.

Simulation environments, synthetic data, ordinary human videos, wearable devices and robots' autonomous exploration are all trying to reduce reliance on real-person teleoperation. The future data structure may be a small amount of high-quality real operations plus large-scale simulation and human videos, instead of requiring unlimited increase in teleoperation duration.

XDOF itself is also expanding into human first-person data and automatic evaluation tools. This shows that the company knows that simply relying on humans to remotely control robots cannot support long-term valuation.

More realistic competition has already emerged. Startups such as Mecka also position themselves as "the data and deployment layer for physical AI", trying to build robot training data from human movements, sensor information and real scenarios.

What this competition will ultimately compete for is not who has the most videos, but who can prove that their data can enable robots to learn tasks faster and work stably in real environments.

V

For China's robotics industry,

the real opportunity may not only lie in "making robot bodies"

The fact that XDOF is sought after by capital also has implications for China's robotics industry.

China already has a dense robot supply chain and a large number of potential application scenarios. In e-commerce warehouses, manufacturing factories, catering stores, property services and household environments, there are massive amounts of real tasks that robots can learn from.

But "having scenarios" does not equal "owning data assets".

If the operation processes in factories are not collected in a standardized way, if the data formats used by different robots are incompatible, if the data only stays in a single project, then even the richest scenarios will hardly form model capabilities that can be reused.

Chinese enterprises currently generally attach importance to robot bodies, joints, reducers and motion control, but may underestimate the value of the data production system.

More important questions in the future include: who is responsible for defining tasks, who collects failure processes, who owns factory operation data, whether data can be migrated across robots, and how this data can be continuously fed back to model training.

For robot manufacturers, selling a piece of equipment is just the beginning. What can truly create a gap is whether the equipment can continuously generate data, improve the model after entering the real scenario, and then deploy the improved capabilities to more robots.

For startups, the opportunity does not necessarily lie in replicating an XDOF.

A more realistic direction in China may be to build a vertical data system around advantageous industries: collect assembly and quality inspection data for automobile factories, sortation and packaging data for e-commerce warehouses, and food preparation and cleaning data for catering scenarios.

The data does not need to cover all robots, as long as it can prove the model success rate and deployment efficiency in a high-value scenario, it can generate commercial value.

At the same time, data collection will also involve employee privacy, labor process records, trade secrets and cross-border transmission. If cameras record how workers work for a long time, who owns these motion data, whether employees are aware of it, and whether the data can be used to train robots for other clients, all need to establish rules in advance.

Conclusion

The rapid rise of XDOF presents a phenomenon with a distinct "vanity fair" flavor in the robotics wave.

In the center of the stage are humanoid robots that can run and dance, but behind the stage are hundreds of operators constantly repeating grasping, folding, wiping and sorting, breaking down the movements that humans take for granted into data that machines can understand.

When everyone is looking for the foundation model of the robotics era, capital begins to realize that models are not the only scarce resource, and organizations that can produce physical world experiences on a large scale are