They raised 17 billion yuan in financing within half a year, and the data "shovel seller" has become the most lucrative business in the embodied intelligence track.
At the end of July, LiberAI completed its Pre-A+ round of financing with a sum of several hundred million yuan. The list of investors includes 360 Group, China Development Bank Capital, Sinovation Ventures, CMC Capital, Bifu Capital and HSG.
This company, actually founded on December 8, 2025, is named "Jiangxian Technology". The implication of its name is that "the desire to rest is the primary productivity"; its registered capital is 1.46 million yuan, and the team has less than 30 members.
IT Juzi has found that this is already its fifth round of financing since its establishment. Over the past more than half a year, it has been advancing financing at a fast pace, completing four consecutive rounds from seed round, angel round, angel+ round to Pre-A round in one go. According to reports, its latest valuation has reached 5 billion yuan.
The founder Liu Songming is a post-2000s generation, winner of the Special Scholarship of the Department of Computer Science of Tsinghua University, top of his major, studied under Professor Zhu Jun, and started his business right after graduating with a doctorate at the age of 23. The RDT-1B he previously led is the first large-scale diffusion Transformer foundation model specially designed for dual-arm robot operation, with a parameter count of 1.2 billion.
What LiberAI targets is human UMI data + world model — to put it plainly, what it sells is the data fed to robots and the "brain" that processes the data.
According to statistics from IT Juzi, in the first half of 2026, 25 Chinese embodied intelligence data startups received a total of more than 17 billion yuan in financing. It can be said that "data" has become the most certain business in the embodied track (excluding a few companies that develop both embodied data models and robot bodies, such as Xinghaitu and Qianxun Intelligence).
To obtain all the companies in the embodied data track and their financing data, you can view and download the data on the IT Juzi album page: https://www.itjuzi.com/album/766 (42 companies are currently included).
Among them, Guanglun Intelligence not only has large-scale external financing, but also won orders of 550 million yuan in Q1.
While gold prospectors have not yet dug up gold, those who sell shovels have made money first.
I. Unsolved Problem: The Data Desert of the Robotics Industry
Why can large language models explode? Because there is inexhaustible corpus on the Internet.
But training robots does not rely on text and language, but on 3D physical motion data such as grasping, placing, walking, grabbing, twisting, and flipping a wok — this type of data hardly exists on the Internet.
The scarcity of embodied data is a major bottleneck across the entire industry.
Grand View Research predicts that the global data collection and labeling market will reach 17.1 billion U.S. dollars by 2030.
At present, the widely recognized structure of embodied data in the industry presents a pyramid shape: the top of the pyramid is real machine data, which has reliable quality and can be directly deployed, but it is also the most expensive; the middle layer is simulation synthetic data + UMI data, with moderate cost and unlimited mass production capacity, but there is a Sim-to-Real gap (approximate deviation of friction/deformation); the bottom of the pyramid is Internet video and human behavior data, which has the widest source, but lacks interaction details and has sparse information.
Then there is a group of companies dedicated to making breakthroughs at each layer — this is a coordinate system for understanding the current entrepreneurship boom in the embodied data track.
II. Panorama of Five Schools with More Than 30 Companies
School 1: Real Machine Teleoperation/Haptic School — Treat Data Production as an Assembly Line
This school follows the most simple logic: build factories, place robots, deploy motion capture equipment, and mass-produce high-precision data through assembly lines.
The most radical one is Pacini Perception.
The company's haptic sensor shipments rank first in the world. In 2025, Pacini built the world's largest embodied data collection factory Super EID Factory in Tianjin, with an annual output of 200 million pieces of data. After completing the B round of financing of more than 1 billion yuan and its valuation exceeding 10 billion yuan in March, it announced that it would build four more super data factories in Suqian, Wuhan, Zigong and Ganzhou.
Veterans in the motion capture industry form another branch of this line.
Dai Ruoli, the founder of Noitom Robotics, once helped Noitom Technology take up 70% of the global professional motion capture market share, and then founded the new company Noitom Robotics. In June, it completed the Pre-A++ round of financing of several hundred million yuan with a valuation of about 3 billion yuan. Beijing and Shanghai AI Industrial Funds, Shenzhen Venture Capital, CICC and Kunlun Capital all participated in the investment, and it released the ModalityNet multimodal data platform.
Qingtong Vision takes the path of open source to build its ecosystem. In May, it released the world's largest 1000-hour high-precision human optical motion capture open source dataset, with an expected annual output of 500,000 hours. The martial arts data collection for the debut robot martial arts performance "WuBOT" of Unitree Robotics at the Spring Festival Gala was completed by its Project Decode system.
Emerging haptic startups such as Lingsheng Technology are also seizing market positions.
School 2: Simulation Synthesis School — "Printing" Data in the Virtual World
If real data is too expensive, can we create a virtual world with sufficiently realistic physical rules in the computer to "print" training data in batches?
This route has produced the world's first unicorn in the embodied data track: Guanglun Intelligence.
In March this year, it completed the A++ round of financing of 1 billion yuan, then received investment led by Ant Group in May, and completed another 1 billion yuan of financing in June. Its post-investment valuation exceeded 2 billion U.S. dollars (about 15 billion yuan).
The revenue of Guanglun Intelligence increased 10 times in 2025, and its new orders in the first quarter of 2026 reached 550 million yuan, exceeding the total amount of last year in just one quarter.
Kuaivi Intelligence is the representative of "generative simulation": its self-developed DexVerse simulation engine automatically generates massive scene data in the virtual world. In July this year, it officially announced 1 billion yuan of financing with a valuation of over 10 billion yuan and launched the IPO process. Its revenue in the first half of the year reached 100 million yuan, and more than 1500 sets of models have been delivered for commercial use.
Among listed companies, Qunhe Technology, the "first global spatial intelligence stock" that was listed on the Hong Kong Stock Exchange in April this year, is extending the massive 3D scene data accumulated from 3D cloud design into the embodied intelligence field.
Simying Technology, a simulation platform, completed a total of several hundred million yuan of financing in the Pre-A and Pre-A+ rounds in February and July, and jointly built a synthetic dataset with the national-local co-built humanoid robot innovation center.
School 3: UMI No-ontology/Portable Collection School — Bypass the Ontology and "Replicate" Humans
Collecting data by remotely controlling robots has a flaw: the collected data is not the real ability of human beings, but the actions compromised to make robots keep up. Therefore, this school simply bypasses the ontology, and lets people work directly by wearing gloves, grippers and wearable devices.
This is the route with the highest capital density this year.
On June 1, Jianzhi Robotics officially announced that it has completed several consecutive rounds of financing totaling hundreds of millions of yuan, led by Ant Group, Didi and Delian Capital, with participation from Shunwei Capital and BV Baidu Venture Capital — this is the largest financing in the no-ontology data field so far.
The founder Chen Jianxing was former senior algorithm director of Momenta. The company was established one year ago, has covered tens of thousands of real scenarios, and has reached a strategic cooperation with Ant Lingbo.
Tishi Zhihang proposed the human-centric data collection paradigm and launched the SenseHub wearable collection solution. In April, it completed the Pre-A round of financing of 455 million U.S. dollars, setting the record for the largest single financing in China's embodied intelligence field. It was jointly led by GL Ventures, HSG and Meituan, and has launched the "Embodied Data Spark Project" with a target of 100 million hours of data.
The FastUMI system of Luming Robotics from Tsinghua University background has increased the data collection efficiency by three times and reduced the cost to one fifth. It has received consecutive leading investments from Mitsubishi Electric and won orders from leading customers.
In February this year, Unitree Robotics spun off its data business into an independent company Mifeng Technology, which completed the seed round and angel round of several hundred million yuan in only ten days, led by HSG, and plans to achieve a data production capacity of 10 million hours in 2026.
Younger startups are also emerging: Xingyi Technology, incubated from the Department of Computer Science of Tsinghua University, follows the path of Nvidia EgoScale, focuses on first-person wearable collection with high degrees of freedom and millimeter-level precision, has completed the first round of financing and obtained the Pre-A round of financing in July. Zizai Wujie founded by Professor Lu Zongqing from Peking University directly uses Internet videos to pre-train general action models.
Qiongche Intelligence, incubated by Flexiv, has obtained nearly 100 sets of orders with its "production-accompanying" data collection system CoMiner, received several hundred million yuan of financing in June, and is about to release its self-developed world model.
There is also a very new company OriginFlow, which adopts the NeuroScale paradigm, captures the electrical signals of human muscle contraction through its self-developed myoelectric collection kit, encodes and reconstructs hand posture/force exertion/haptic feedback through the PULSE foundation model, connects the native conduction link of "intention — muscle — action", and makes up for the limitations of UMI such as visual occlusion failure and lack of force and haptic perception.
Founded one year ago, the company has completed a total of more than 500 million yuan of financing in the angel round, strategic round and Pre-A1 round. Its founder Qin Shentao is a post-2000s Tsinghua doctor and graduated from Harbin Institute of Technology with a bachelor's degree.
School 4: Video Distillation/World Model School — Teach AI to "Understand" the Physical World
This route solves the problem of "turning waste into treasure" at the bottom of the data pyramid: there is inexhaustible human operation videos on the Internet — cooking, handicrafts, maintenance, but these videos lack the force perception, haptic perception and 3D trajectory information required by robots, and have always been dead data that "can be seen but cannot be used".
The idea of the video distillation school is to use algorithms to reversely deduce action trajectories and physical interaction information from 2D videos, distill the dead data into trainable embodied data, and the cost can be reduced to a few thousandths of that of real machine collection.
The world model school is more straightforward: first let AI learn the operation rules of the physical world, and then generate unlimited training data "in its mind".
The SynaData pipeline of Shutu Technology can batch extract multimodal embodied data such as hand trajectories and object motion paths from Internet videos, with extremely low comprehensive data collection cost. It has received tens of millions of yuan of investment led by Oriental Fortune Capital, and its data is adopted by mainstream open source models such as Tsinghua RDT and Unitree UniVLA.
Jiaji Shijie uses the world model GigaWorld to generate training data, which improves the performance of the VLA model by nearly 300% in three generalization dimensions. It completed the B+ round of financing of 1 billion yuan in June, with a total financing of 3.5 billion yuan in 3 months.
Shendu Jizhi also bets on human first-person video data. It completed three rounds of financing totaling hundreds of millions of yuan+ in the first half of this year. In May, the Z-WM world model of Shendu Jizhi ranked first in the global ranking in the WorldArena evaluation.
LiberAI mentioned at the beginning of this article is also a hybrid of this route and the world model track.
School 5: Data Infrastructure/Platform School — It Is Both a Data Platform and an "Oil Refinery"
The consensus of this school is: the bottleneck of embodied data lies not only in "how to collect", but also in "how to use".
Different robot bodies and different sensors of various companies lead to a wide variety of data formats. The raw data collected is like crude oil with mixed components. If it is directly fed into the model, the model performance will get worse and worse during training. Who will do the work of cleaning, aligning, labeling and evaluating to refine crude oil into standard gasoline?
What's more, training grounds, collection equipment and motion capture systems are all heavy assets that small and medium-sized teams cannot afford. Therefore, it is necessary to build shared training grounds, set data standards, and sell toolchains and platform services. It does not bet on which data technology route will win, but bets that no matter who wins, they have to pass through the path it has built.
Wuwen Zhike relies on the Deqing data collection training ground with virtual-real integration closed loop, producing thousands of hours of data per day. It completed more than 100 million yuan of financing in April, and signed hundreds of millions of yuan of orders with ByteDance, Wujie Power and other companies in the first quarter.
Yiren Technology completed two consecutive rounds of 100-million-level financing in April, and announced that its revenue in 2025 exceeded 100 million yuan and achieved profitability. It may be the first company in the track to announce that it is making profits.
There is also Zhiyu Jishi founded last December, which focuses on the midstream transformation of data cleaning, alignment and governance. All ontology companies such as Lingchu Intelligence, Qiongche and Zhi Pingfang have invested in this "upstream" company.
Listed companies have also entered the market. Haitian Renshi, a leading data service provider, cooperated with the Beijing Shijingshan Humanoid Robot Data Training Center to jointly build the "embodied intelligence data training ground".
Large manufacturers are also getting involved. JD released the full-link infrastructure for embodied data, planning to mobilize 600,000 couriers and riders to collect data through crowdsourcing, with the goal of accumulating 10 million hours of real scene videos in two years; Baidu is building a "data supermarket".
III. Calm Down: Is the Demand Side of Data Rigid?
2026 has become the first year of large-scale development of embodied data. According to statistics, as of the end of April 2026, there are at least 90 data collection centers in the status of "in use or planned for construction"; among them, 64 have been put into use, and the rest are under construction or in planning.
But the hotter the trend is, the more we need to pour some cold water on it.
The embodied data business has a structural problem that is rarely pointed out: its demand side and supply side are burning money from the same group of capitals.
Looking at the buyer list, you will understand — large model teams, startup robotics companies, and transforming ontology manufacturers, most of these data purchasers are not profitable yet, and their procurement budgets come from the financing funds they just received.
In other words, the revenue of data companies is essentially the secondary distribution of downstream financing: embodied intelligent robot companies get financing, then spend money to buy data to tell a good story, and then raise the next round of financing with the good story.
If the landing progress of robot ontology companies fails to meet expectations, data orders and financing ebb will disappear at