To stockpile data for its robots, Figure is offering a global bounty to recruit humans to "do the work".
Brett Adock, Founder and CEO of Figure AI, the image is processed by AI
Humanoid robots need to cross a critical threshold before they can truly step into the real world: data.
Over the past few years, AI robotics company Figure has been working to build a general-purpose humanoid robot intelligent system. It launched Helix, a foundational model for robots, aiming to endow robots with human-like capabilities in visual understanding, task planning and motion execution, gradually moving them from laboratories to homes, factories and commercial settings.
However, as the model's capabilities continue to improve, an increasingly obvious problem has emerged: robots lack sufficient real-world experience.
For large language models, the internet provides massive amounts of text as learning materials. But robots operate in the physical world, and need to understand the complex interactive relationships between humans, objects and the environment. These capabilities cannot be obtained solely from text and simulation environments, but require a large volume of real-world human behavior data.
On August 25 local time in the United States, Figure launched the Index project, which was simultaneously released on iOS and Android platforms, hoping to collect real task data from global users to build a training data system covering the real world for robots.
Up to now, Index has covered 108 countries and regions around the world, with a total of more than 264,000 app downloads, over 44,000 weekly active users participating in data contribution, and more than 16 million videos uploaded.
Figure AI stated that it will invest more than 1 billion US dollars in data and computing resources in the next 12 months.
Index has more than 44,000 weekly active users
01. From "Able to Act" to "Able to Understand the World", Robots Need Human Experience
Founded in 2022, Figure has always focused on advancing the R&D of foundational robot models, one of the core directions of which is Helix, the Vision-Language-Action (VLA) model.
This type of model is designed to free robots from reliance on fixed programs, enabling them to observe the environment, understand tasks, and continuously complete complex actions.
However, the data required for robot model training is completely different from that of traditional AI. Ordinary videos can only tell the model "what happened", but what robots need to understand is: why the action is taken, and what to do next.
For example, when a person tidies up a desk, it seems to be just moving items simply, but it contains a large amount of implicit information: how to judge the category of items, how to plan the placement position, and how to adjust actions according to spatial changes.
These experiences hidden in daily behaviors are exactly the key for robots to achieve generalization capabilities.
Previously, Figure tried to obtain training data through third-party data suppliers, but the company found that the externally supplied data could not meet the development needs of Helix in terms of scale, diversity and quality.
Therefore, Figure AI decided to build its own data collection network, and Index came into being.
The Index mobile app shows household tasks that creators can record and upload
The project invites ordinary users to become "Creators", who record task processes in the real world through mobile phones or head-mounted devices, and these data are used for robot training.
The scope of tasks is not limited to standardized operations. Users can record household activities such as tidying up rooms, folding clothes, and cleaning kitchens, as well as work processes in logistics centers, restaurants, factories and offices. Even some low-frequency but valuable tasks, such as cleaning cat litter and changing engine oil, will become data sources.
What Figure hopes to collect is not just simple action videos, but also various "unexpected situations" that robots may encounter in the future. Because after robots truly enter the real environment, the biggest challenge may not be completing actions they have learned, but making correct judgments when facing unfamiliar scenarios.
Figure stated that for every 1000 hours of data collected, Index data covers an average of 373 unique tasks, 1146 different operating objects, and 116 different environments.
Everyone can become a creator of Index, and they can also book creators through the app to come to their home to assist with housework or commercial tasks
These complex samples from the real world are exactly the foundation for robots to evolve from "being able to perform actions" to "understanding the world".
02. 16 Million Videos, Turned into "Fuel" That Robots Can Learn From
But collecting data is only the first step. For robot models, a large number of unprocessed human videos do not equal effective training data.
Real-world data often has a large number of problems: some people upload low-quality videos, some repeatedly submit similar content, and some deliberately create invalid data in order to obtain rewards.
Therefore, Figure has built a data processing system to convert original videos into data assets that robots can understand and learn from.
Five steps for Index to process the data submitted by creators
First, the system needs to complete data screening. All uploaded content will go through automatic detection to judge whether the video is clear, the task is complete, and the content meets the training requirements. Data that cannot reflect effective human-computer interaction will be filtered out directly.
Subsequently, Figure will conduct authenticity verification. Since creators can get paid according to their contributions, the company needs to identify abnormal behaviors to prevent a large amount of low-value or even false data from entering the training process.
On this basis, the system will also deduplicate and reorganize the data. For robot training, more data is not always better. If a large number of similar tasks appear repeatedly, the model is prone to deviation. Therefore, Figure AI will identify highly similar videos through algorithms, and adjust the data proportion among different tasks, environments and objects.
Finally, the screened data will be further labeled, so that the model can understand the task objectives, action processes and environmental changes.
In other words, the core value of Index is to build a complete link from real behavior collection, data processing to model training, enabling human experience to be continuously transformed into the intelligent capabilities of robots.
At present, Index can process about 30 minutes of video uploads per second, which is equivalent to receiving about 4.9 years of human working time data every day. At the same time, Figure has paid more than 15 million US dollars in remuneration to creators.
03. Data is Becoming a New Entry Point for the Robotics Industry
The goal of Figure launching Index is not only to train the current robot model. The deeper significance is that the company hopes to build the data infrastructure for the robotics era.
In the past, the development of visual AI was inseparable from datasets such as ImageNet. The breakthrough of large language models was built on the scale of internet text. The future competition in robotics may depend on who owns the largest, most authentic and richest real-world interaction data.
Figure is trying to build a similar infrastructure.
At present, the company has promoted humanoid robots into some industrial scenarios, including cooperating with enterprises such as BMW, and plans to promote the large-scale production of Figure 03.
However, Figure AI believes that hardware manufacturing is not the biggest bottleneck for the commercialization of robots. What truly determines whether robots can enter more scenarios is whether robots are intelligent enough to handle complex tasks in the real environment.
In the future, the data collected by Index may not only be used to train Figure's own robots, but also become part of the robot service system.
Today, humans use platforms to find service personnel to complete housework and work. Figure hopes that these tasks can be completed by robots in the future. In this process, whoever masters the real-world data will likely hold the entry point of the next-generation robotics industry.
This article is from the WeChat Official Account "Tencent Tech", author: Zhou Xiaoyan, published with authorization from 36Kr.