HomeArticle

Embodied Intelligence is recruiting "data riders"

全天候科技2026-09-24 17:58
Be a teacher for robots

On September 23, Mifeng Technology officially launched its crowdsourced data platform "Mifeng Pai", attempting to massively open up embodied intelligence data collection to the general public.

Founded in February 2026, Mifeng Technology's business covers the collection and processing of real-device data, robot-free data and simulation data, positioning itself as a third-party data platform serving the entire industry.

With the help of the "Mifeng Pai" platform, ordinary people can apply for MEgo devices, accept collection tasks through the App, and the platform will settle remuneration based on the verified duration of valid data after review. 

Mifeng is also opening recruitment for "robot trainers", and has jointly initiated a scenario data alliance with more than 50 enterprises and institutions in the fields of hotels, supermarkets, manufacturing and other sectors. The company aims to disperse the data production that was originally concentrated in data collection factories into more real work and life scenarios.

Over the past year, the answer to "where data comes from" in the embodied intelligence industry has been evolving rapidly. The most straightforward solution in the early stage was to centralize robots in data collection centers, where operators remotely control real robots to repeatedly complete tasks such as grasping, transporting, and sorting.

However, real robots are expensive, with limited collection efficiency, and it is very difficult to push the data scale from millions of hours to tens of millions or even hundreds of millions of hours by simply increasing the number of robots and collectors linearly.

As a result, the industry has begun to decouple more data production processes from the robot body itself.

Mifeng has chosen to further open up the production end: allowing more people to enter real scenarios with lightweight devices, and the platform will uniformly process these decentrally generated data into training data that models can use.

When embodied data begins to try crowdsourced production, will it become a rapidly expandable data network, or a more complex labor business?

The data version of Meituan: how to calculate the accounts?

For ordinary people, participating in Mifeng Pai generally needs to go through several steps: download the App, rent MEgo devices, receive tasks, complete specified operations in life or work scenarios, upload the data, obtain commission according to the duration of valid data and withdraw cash.

Users can select tasks in the task hall, and ongoing tasks adopt a duration quota system. The rent of MEgo device is 39 yuan per day, and the discount price during the promotion period is 19 yuan per day. The basic return of tasks displayed on the App is about 20 yuan per valid hour.

This is not the unified hourly salary that collectors finally get.

According to the company, the remuneration consists of the basic urban pricing and dynamic subsidies: the former refers to the income level of different regions, and the latter depends on the task completion efficiency and data quality.

Tasks range from a few minutes to more than an hour, and are finally settled according to the verified valid duration.

After a certain type of data is fully collected, the platform can stop distributing tasks; for scenarios that are scarce or difficult to access, additional subsidies will be used to improve attractiveness.

The currently open collection tasks cover more than 20 fields including production maintenance, catering services, logistics handling, elderly care and housework sorting. Daily operations such as clearing badminton buckets in badminton halls and replenishing drinking cups in convenience service stations are also included in the collection scope.

"Crowdsourcing" is easily associated with a variant of the food delivery rider model: the platform connects demands with scattered workers, and then schedules them through task rules and prices.

But the delivery difficulty of the two is not the same.

Food delivery only needs to deliver the goods to the designated location, while embodied data must meet the specific requirements of the model for scenarios, actions and quality, and scattered collectors have great differences in operation habits, professional skills and working environment.

From the perspective of business model, Mifeng believes that crowdsourcing can first reduce the labor cost at the collection end.

Traditional centralized data collection requires specially hiring collectors to repeatedly complete tasks in a fixed site.

The crowdsourcing model hopes to embed data collection into the original life and work of participants. The platform pays for the incremental data contribution, rather than the full cost of hiring a full-time collector.

Accordingly, the complexity of back-end operations will also increase.

The demands of model companies first need to be decomposed into tasks that ordinary people can understand, with clear action steps, equipment positions, completion standards and acceptance rules.

The platform also needs to complete basic training, equipment collection and transfer through offline service stations, so as to minimize the operation deviation among different collectors.

Unified devices solve part of the input differences. The data that passes the initial screening will continue to go through action segmentation, labeling, trajectory extraction and manual sampling inspection before it can become data usable for the model.

Mifeng does not simply divide data into "usable" and "unusable", but grades it according to quality and use value.

Data with obvious shooting failures will be eliminated, and the rest of the data may be used for pre-training, specific task training, model evaluation or long-tail scenario learning respectively.

Yao Maoqing, Chairman and CEO of Mifeng Technology, said that more than 95% of the data that passed the initial settlement screening can still be utilized after post-processing.

According to the data disclosed by Mifeng, the number of registered users of Mifeng Pai reached 20,000 within one month of internal test, with a total of 13,000 task submissions, and the participant with the highest income earned more than 5,000 yuan in a single month.

The more scattered the scenarios, the more complex the operation

When Mifeng launched the crowdsourcing platform at this point in time, there were two simultaneous changes in downstream data demands: customers need larger and larger data volume, and the required scenarios are becoming more and more detailed.

Yao Maoqing said that millions of hours are changing from the industry supply target to the training demand of a single customer. Some customers have put forward demands at the level of tens of millions of hours, and some leading customers even require suppliers to deliver 50,000 to 100,000 hours of data per week.

However, if the same type of data is only produced ten times more, the value of crowdsourcing is still limited.

In the past, embodied data collection was concentrated in family scenarios, and a large number of teams repeatedly collected tasks such as folding clothes, and this type of data has become saturated. What is becoming scarcer are new skills, new environments and new operation processes that the model has not seen before.

Mifeng is actually expanding scenarios on both the demand side and the supply side.

Mifeng Technology said that the company will regularly hold production plan review meetings to summarize market orders, first judge which demands can be met by existing inventory, and then allocate the missing parts to different collection paths according to order value and resource conditions.

Yao Maoqing also mentioned that leading customers have formed some common requirements in real scenarios, spatial accuracy and action semantic labels, but specific demands will continue to sink to different processes and actions.

Centralized data collection centers can replicate standard environments such as kitchens, living rooms and desktops, but it is difficult to reproduce scenarios such as automobile maintenance, hotel services, agricultural production or specific factory production lines in a fixed site.

Relying on a full-time team to find these scenarios one by one is costly and slow. Crowdsourcing disperses scenario finding and data collection to people who are already working in different industries, making it easier to reach long-tail professional scenarios and their operation skills.

At the same time, Mifeng Technology hopes to connect scenario parties such as factories, communities, hotels, stores and logistics centers in advance through the "Scenario Data Alliance", and solve problems such as security, authorization and cooperation before customer demands arrive.

Zhang Zhifu, Vice President of Mifeng Pai Business, also mentioned that in the future, the platform may match tasks according to the occupation and historical tasks of collectors, and open up independent scenario declaration, and the platform will review and judge whether it matches downstream demands.

In his explanation, it is difficult for the internal team of the platform alone to list thousands of tasks worth collecting in advance. People who are actually on the front line of agriculture, maintenance and production know better what scenarios and skills they have.

What Mifeng ultimately wants is not a larger data collection factory, but a scenario network that can be scheduled according to orders.

This also makes operation one of the core capabilities of embodied data companies.

Yao Maoqing judged that when the industry truly enters the stage of large-scale development and profitability, it will "attach great importance to operational capabilities":

It is necessary to convert the data gap of model companies into tasks, find the corresponding people and scenarios, process highly heterogeneous data into unified specifications, and maintain cost advantages while delivering on schedule.

More importantly, with the expansion of procurement volume, customers' requirements for supplier scale, quality and delivery capability are also increasing. Based on this, Yao Maoqing judges that embodied data suppliers may further concentrate, and even show an oligopolistic trend.

Mifeng said that supporting the production of tens of millions of hours of data requires investing hundreds of millions of yuan in fixed assets such as collection equipment, personnel settlement, data storage and computing power infrastructure.

In the future, the industrial chain may see further division of labor in hardware, operation and data post-processing, but end-to-end coverage of the entire link is not a business that ordinary entrepreneurial teams can easily complete.

In addition, for model companies that need millions or even tens of millions of hours of data, managing dozens of suppliers at the same time means that they have to deal with the differences brought by different hardware, formats and quality systems by themselves.

The more suppliers there are, the higher the cost of data cleaning, conversion and training verification will be.

In other words, the scenarios can be scattered, but the delivery interface is better to be centralized.

Whether it is expanding the data volume or opening up scenarios, the ultimate competition is not only for current orders, but also for occupying the ecological niche of data production and scenario entrance in advance before the demands continue to change.

Who will pay for the next 100 million hours of data?

The Ego robot-free collected data represented by Mifeng Pai crowdsourcing currently mainly serves the pre-training of embodied intelligence companies, world model companies and general large model companies.

The model learns the relationship between objects, actions and environments from human operations, so as to expand its coverage of different scenarios and skills.

In simple terms, robot-free data solves the problem of whether the model "sees enough", while real robot data solves the problem of whether the robot "can actually complete the action".

Yao Maoqing said that if you only need to complete a specific task Demo, using the target robot to collect real machine data will have a more direct effect; when the sample size of Ego data is small, its improvement on a single task is not obvious, and its value is more reflected in the Zero-shot generalization ability when facing new tasks and new environments after the scale expands.

As models begin to make decisions in the physical world, what needs to be supplemented is not just more hours of data, but also richer scenarios, higher spatial accuracy and more complete action semantics. The stronger the model capability, the more complex tasks it can learn and enter, and new data gaps may also appear accordingly.

At least under the current training paradigm, robot-free data is not a temporary substitute when there is a shortage of robots.

A longer-cycle variable is that robots themselves begin to produce data.

For mature tasks such as box moving, sorting and fixed production line operations, running robots can directly record joint status, control instructions, as well as success, failure and abnormal situations.

These data are naturally consistent with the target robot body. With the large-scale deployment of robots, the relative importance of human collection in mature tasks will decline.

For Mifeng, the more critical question is whether customers will continue to purchase from third-party platforms for a long time even if the industry still needs a large amount of data.

Once leading robot companies form their own installed scale, scenario network and data return system, the economy of self-collection will improve, and the data production right of mature tasks may also shift to robot operators.

The fact that human data still has value does not automatically mean that third-party data collection platforms still have value.

Yao Maoqing positions Mifeng as a "turnkey" data service provider.

He believes that most embodied companies still tend to concentrate resources on hardware body and model algorithm R&D, and entrust heavy asset and heavy operation links such as equipment, manpower, scenarios and data operation to professional platforms. Part of the standardized data can also be sold across customers for many times to reduce repeated collection of similar data. 

How far human crowdsourcing can go ultimately depends on two things: how many worlds robots have not yet entered, and whether third-party platforms can find and produce these data at a lower cost than customers themselves.

This article is from the WeChat official account "All-Weather Tech" (ID: iawtmt), written by Liu Yichen, published with authorization from 36Kr.