HomeArticle

Data Acquisition Center: A Driving School for Robots, or a Besieged City of Capital

AI前线2026-09-21 09:40
Who is paying for the demand?

In less than two years, at least 90 embodied intelligence data collection centers have sprung up across China.

One in Shijingshan, Beijing, covers an area of 10,000 square meters, is equipped with 120 robots, and has generated a total of 1.17 million pieces of operation data. Paxini's center in Tianjin spans 12,000 square meters, claims to produce 200 million high-quality training data entries per year, and has built four new centers in Suqian, Wuhan, Zigong, and Ganzhou respectively. ZhiYuan has deployed more than 100 robots in Pudong, Shanghai. JD has also entered the sector, claiming to build the "world's largest" center, mobilizing 600,000 people for data collection with a two-year target of 10 million hours of data.

Local governments are the main driving force behind the mushrooming of data collection centers.

From first-tier to third- and fourth-tier cities, from coastal to inland regions, data collection centers have become a new standard for industrial park investment attraction and a new calling card for future industries. After the decline of land transfer revenue, attracting robot enterprises to settle down with the narrative of data assets has become a new idea for many local governments.

However, beneath the prosperity, doubts erupted intensively in the second half of 2026.

In August, the Financial Times published "Who Is Buying China's Robots", labeling this model with a highly damaging term — circular financing. It means that robot enterprises sell robots to data collection centers built by local governments, and then the funds flow back in the name of data procurement. The book revenue goes up, the valuation is raised, but the real buyer of the data is still the enterprises themselves.

In September, Shao Tianlan, founder of Mech-Mind, posted two consecutive Moments, directly pointing out that "gathering-type" enterprises in the industry create false revenue through data collection centers and related transactions, and explicitly named Galaxy Universal, which has a valuation of over 20 billion yuan and is sprinting for IPO.

The attitude of the regulatory authorities is also tightening. According to The Information, the China Securities Regulatory Commission has issued window guidance to some investment banks and investment institutions: the IPO application of a humanoid robot enterprise to be listed may only be considered if it can prove that it has sustainable revenue, is moving towards loss reduction, or has real technological innovation.

The capital market reacted even earlier. Unitree, the first listed humanoid robot stock, hit 1100 yuan per share with a market value of 444.9 billion yuan on its first trading day on August 19. A month later, its price fell below 500 yuan, with a market value evaporation of more than 240 billion yuan, almost halved.

The core of the controversy is inseparable from the term "data collection center".

Is it a key infrastructure for embodied intelligence to break through the data bottleneck, or a capital game for enterprises to raise valuations and local governments to meet investment attraction targets? Is it another innovative practice of China's hard technology industrial policy, or a repackaged form of circular procurement?

The 95% Supply Gap

First, let's talk about why embodied intelligence needs data collection centers.

The answer is a shortage of data, and an extremely severe shortage. Large models feed on information from the digital world, while robots need interactive data from the physical world.

For example, if you tell a large language model "pass a glass of water", it can understand perfectly, but if you let a humanoid robot actually pick up a glass of water, it will most likely fail. Details in the physical world such as the friction force of the cup, how much force the gripper needs to apply to avoid breaking it, and what acceleration to maintain when walking so that the water does not spill, do not exist on the Internet.

Zhu Kai, General Manager of the Shijingshan Training Center, summed up this pain point in six words: "Good at dialogue, bad at practical operation".

According to the White Paper on Embodied Intelligence Data Industry Research, embodied data has a pyramid structure: the bottom layer is Internet video and synthetic data, which is large in volume and low in cost but has the lowest accuracy. The middle layer is motion capture data, where people wear motion capture suits to perform actions for robots to learn, with improved accuracy but still separated from real physical interaction. The top layer is teleoperation data, where people directly control robots to operate, with force, tactile, visual and joint angle all synchronized, which has the highest accuracy but is also the most expensive.

What data collection centers do is exactly the work at the top of the pyramid.

How expensive is it? Staff from Paxini did the math: a single robot body costs 600,000 to 700,000 yuan, data collectors need to be a professional team with industrial experience, a large site starts at 10,000 square meters, plus hundreds of PB of storage and computing power. All in all, the investment in a standard data collection factory is by no means a small amount.

What's more critical is the scale gap. Guojin Securities estimates that to train an embodied intelligence model with practical application capabilities, at least 10 million hours of multi-modal interactive data are required. However, the total global accumulated data volume is less than 5% of the demand.

A 5% supply means a 95% gap.

The reason why this gap is so hard to fill is that simply piling up equipment is far from enough. Based on the global average level, the labor cost of a set of collection equipment plus the salary of collection personnel is already extremely high. When multiplied by 10 million hours, the scale of investment cannot be achieved just by building a few more data collection centers and adding a few more workstations.

More data does not always mean better. In the early stage, breadth is needed, but in the later stage, depth is required. For example, robots for Chinese families must be familiar with the space, items and habits of Chinese households. The data closer to commercial scenarios is harder to obtain and more expensive.

This is the fundamental landscape of the whole industry: data is extremely scarce. Whoever can accumulate enough data first is likely to take the lead in model capabilities. Precisely because of this scarcity, the business around data has become extremely active and extremely complex.

The National Race for Data Collection Centers

Before 2025, "data collection center" was still a niche concept in the industry. The turning point came in the second half of 2025, after which the development became unstoppable.

March 2026 was a month of concentrated outbreaks. The Southwest Embodied Intelligence Industrial Base co-built by ZhiYuan and Pidu District, Chengdu, started trial assembly. The humanoid robot data training center co-built by Leju and Pingyin County, Jinan, went online. The third phase of the Shijingshan Training Center in Beijing was launched. Paxini announced the construction of four more factories in Suqian, Wuhan, Zigong and Ganzhou to build a "nationwide distributed collection matrix".

According to statistics from Interact Analysis, by the end of April 2026, at least 90 humanoid robot data collection and training centers across the country had been put into use or under construction, of which 64 were already in operation. 15 of them are large-scale, and second- and third-tier cities such as Zhengzhou, Wuxi, Jinan and Mianyang are also following up, with a floor area of mostly 3,000 to 5,000 square meters, and some reaching 10,000 square meters.

Why are local governments so active?

At the policy level, embodied intelligence was included in the government work report for the second time in 2026, listed as a key direction for future industries. Once the top-level tone is set, local authorities will naturally follow up.

A more practical motivation is investment attraction. The settlement of a leading robot enterprise can drive component suppliers in the upstream, promote application scenario development in the downstream, and attract a large number of high-end talents in the middle. For local governments that are looking for new economic growth points, especially those facing declining land transfer revenue, data collection centers are a perfect starting point — investment, employment, industrial chain and technological reputation are all covered.

In the past, attracting manufacturing industries by transferring industrial land at low prices earned land transfer fees and taxes. Now, attracting robot enterprises by investing in the construction of data collection centers uses the same land, but the buildings are changed from factories to training venues, which some people call "data finance".

But whether "data finance" can be established depends on a core question: is the collected data really valuable?

If the data has a market, has buyers, can feed back the model and generate real value, then the data collection center is a data mine. But if the only buyer of the data is the robot enterprise that generated the data itself, then "data finance" is not revenue, but repackaged expenditure.

In August 2026, the FT report labeled China's embodied intelligence industry with the term "circular financing".

This term was originally used to describe NVIDIA: NVIDIA invests in AI startups, the startups use the money to buy NVIDIA GPUs, the funds flow back after a cycle, and both revenue and valuation expand at the same time.

But the Chinese version has more far-reaching implications, because NVIDIA's cycle occurs between enterprises and private capital, which is a commercial behavior, while China's cycle involves local governments and public funds.

As for how this cycle works, let's walk through the process together.

The first step is to build the center. Robot enterprises negotiate cooperation with local governments, the two parties set up a joint venture, the government provides the main funds through state-owned platforms, and the enterprise provides technology and brand.

The second step is to sell robots. The joint-venture data collection center purchases dozens or hundreds of devices from the robot enterprise, with a single unit price of 600,000 to 700,000 yuan. For a data collection center in a second-tier city, the single procurement amount is usually tens of millions to hundreds of millions of yuan, adding a sum of revenue to the enterprise's books.

The third step is to buy data. After the center is put into operation, the enterprise pays the center in the name of "purchasing data services", and the funds flow back.

After a full cycle, the funds provided by the government become the assets and operating costs of the center, and the enterprise's financial statements record a sum of revenue and a sum of data procurement cost. But who is the final consumer of the data? It is the robot enterprise itself.

A staff member of the training center interviewed by FT admitted: Only a small part of the data is sold to buyers other than robot manufacturers.

Independent technology analyst Zhao Saipo summed it up accurately: "The difference between independent demand and demand created within a policy-supported ecosystem has thus become blurred." In other words, when buyers are created within the same system, the word "demand" loses its meaning.

The Dual Nature of Data Collection Centers

At this point, it is easy to have a misunderstanding that data collection centers are scams.

The fact is not that simple. Data collection centers themselves have real value.

The industry does lack data, robots do need practice, and there is nothing wrong with local governments wanting to develop future industries. Data is the core production factor of embodied intelligence — this is not only the statement of enterprises, but also clearly pointed out by the National Development and Reform Commission that it is necessary to build a real-machine data collection system to solve the "data hunger".

A normal data collection center should form a positive closed loop: robots enter real scenarios → generate valid data → data is used for model training → capability improvement → return to scenarios to complete more tasks → customers are willing to pay continuously → new orders and more data.

Every link creates value.

Are there such positive cases? Yes.

For example, the cooperation between Stone Robotics and Aptiv: the R&D team went to the front line of the workshop, spent months refining the process details, the robots expanded from a single workstation to the entire production line, hundreds of robots work stably in clusters, and are moving towards large-scale deployment at the thousand-unit level.

Shanghai Newzhi Embodied takes the tactile route. Spun off from Fudan University, it has built a 1,000-square-meter tactile data collection factory to collect full-dimensional physical information for fine manipulation, and its visual-tactile sensors have entered the verification system of leading customers.

These companies also carry out data collection, cooperate with the government, and obtain financing, but the difference is that the data really goes into the model, the improvement of the model really brings orders, the orders really come from end users, and users are really willing to pay continuously.

What do the deteriorated data collection centers look like? Robots are procured, centers are built, funds are invested, but most of the robots are idle most of the time. The data is collected but not really used for training. Customers are only the government and related parties, with no repurchase and no end-user demand. Funds circulate within the system, the books look good, the valuation goes up, but the robots still cannot work.

The two may not be distinguished at a glance from the financial statements, but three hard indicators can act as a magic mirror to identify them.

The first is the adjusted repurchase rate: after excluding capital-backed customers and related procurement customers, are the real market-oriented end users willing to continue to buy after the pilot?

The second is the real destination of the data. Is the collected data really used for model training? Is there verifiable improvement in model capabilities? Or is the data just stored on the hard disk for storytelling?

The last indicator is operating cash flow. Growth that loses money on every unit sold and relies on financing to sustain is not real growth.

Therefore, data collection centers are not inherently guilty. They can be industrial accelerators or capital treadmills. The key does not lie in the form, but in the content.

Is the Old Path of the New Energy Industry Viable?

People who support the data collection center model love to talk about new energy vehicles and photovoltaics.

The logic is straightforward: in the early days, the new energy industry was supported by government subsidies, and photovoltaics also became the world's number one with policy support. The government first supports the initial demand, and when the industry matures, real demand will naturally follow. This is China's proven successful approach.

This analogy sounds reasonable, but there is a key difference.

The subsidies for new energy vehicles are given to existing real buyers. Public transport companies need buses, taxi drivers need taxis, and consumers want to buy cars. Subsidies reduce the purchase cost, but do not change the fact that "they need this product".

The data collection center model is different. It is not subsidizing existing buyers, but creating a buyer that would not have existed otherwise. Without data collection centers built under government leadership, who would buy these robots? Are real industrial and commercial customers willing to procure them at the current price and scale?

Subsidizing buyers helps the industry cross the initial cost threshold. Creating buyers helps the industry bypass the market verification process. This is the essential difference.

Another reference is the US DARPA model: the Department of Defense funds challenge competitions, universities and enterprises provide solutions, and industrialization depends on their own capabilities. But DARPA's funding is R&D investment, not operating revenue. No one counts DARPA's grants as revenue.

The subtlety of the data collection center model is that it packages R&D investment as operating revenue. R&D should have been recorded on the cost side, but it has become a highlight on the revenue side. The statements look better, the valuation is higher, financing is easier, but no one knows how much technology has been improved and how much data has been accumulated.

When the dotcom bubble burst, the companies that were eliminated were those with no revenue, and the ones that survived were products with real users. But the problem facing China's embodied intelligence industry may be more tricky: it may not have even found real users yet.

Of course, hard technology industrialization requires patience, and requires long-term investment from the government, capital and enterprises. But patient investment and bookkeeping games are two different things. Patient investment means really spending money on R&D, data collection and scenario construction. Although it is slow, every step counts. Bookkeeping games rely on related transactions to make revenue