The world's largest, Google is queuing up to buy Figure, and this team has built the most popular robot in Silicon Valley.
There's no hiding it anymore, not at all this time!
Google DeepMind, Figure, 1X, Genesis AI, Dyna Robotics……
The most popular robotics players in Silicon Valley in this round all turn out to share the same data "supplier" behind the scenes.
It is behind the 1 million-hour Scaling Law that Dyna-2 has achieved.
The first-person data used by GENE-26.5 to complete a 20-step meal in one go also comes from here.
To win these top-tier clients, being only good at "data collection" is obviously far from enough.
This mysterious company has another more hardcore ace in its hand — powerful data processing and algorithm capabilities.
For the most intractable corner cases in the industry, such as scenes where hands are blocked, wearing transparent gloves, or movements are too fast to be clearly captured, they can still extract usable materials for model learning frame by frame.
It is precisely because the entire chain of collection, processing, algorithm and delivery is fully connected, that it has accumulated a high-quality data stock of 1.5 million hours, and completed a total of 2 million hours of data delivery worldwide.
Stable recognition can still be achieved even with fast movement and multiple hands in the frame
Reaching this scale basically means it has become the data infrastructure behind top robotics companies.
Now, this mysterious company is finally unveiled — it is Maxinsights!
A Physical AI team rooted in Silicon Valley, with almost all top clients in the United States, quietly launched its business in China only in March this year.
What it aims to do is to scale human experience in the real world into training fuel for Physical AI.
Human Daily Videos Become Intensive Training Courses for Physical AI
So here comes the question: why does a data company need to train its own model? And why make the model so highly precise?
The answer starts from the threshold for Physical AI to step into the real world.
There are a huge number of details between letting a Physical AI explain how to fold clothes and letting it actually fold clothes well.
Where should the hand grab? What if the corner of the clothes is pressed down? How to adjust the next movement after the fabric slides?
These experiences are rarely fully written down in text, but humans use them every day.
Large language models can implement Scaling Law because the Internet has already accumulated enough text for them; vision models have also inherited billions of ready-made images.
Robots have never had such conditions. There are billions of hours of videos on the Internet, but almost none of them record how people work from the "hand perspective".
First-person data, also called Egocentric Data or Ego Data, provides a recording method —
From the perspective of the operator, shoot how both hands interact with objects, and what changes occur in the environment before and after the movement.
It makes the learning materials closer to the task itself.
In the past, the mainstream way to teach robots to work was teleoperation. That is, a person wears a device to remotely control the robot to grab a cup or fold a towel, and the robot records every movement.
This kind of data has high quality, but it is expensive and slow. One robot can only produce a few hours of data a day, and the data is tied to a specific robot body, which means re-collection is required if you change to another robot.
This is exactly where the value of first-person data lies.
The collection device follows the operator, records the scene he sees, and the interaction process between his hands and objects.
The operating experience accumulated by people in the kitchen, workbench and living space thus gets the opportunity to enter the training system.
However, shooting the video is only the first step.
Video Is Not Equal to Training Data
Take the video of a human unscrewing a bottle cap in the kitchen as an example. A human can tell what is happening at a glance.
After entering the training process, further decomposition is required —
Which hand holds the bottle body, which hand starts to rotate, which frame the movement starts at, which frame it ends at, and how the wrist moves in space.
If two hands are crossed, fingers are blocked by the bottle body, or the movement is too fast to cause blurriness, the judgment will become difficult.
A label like "unscrewing the bottle cap" cannot explain these details.
What robots need is structured data, not just a piece of video.
What exactly makes Maxinsights' algorithm so strong? A demonstration introduced by the team integrates three layers of capabilities into the same piece of data.
What Problems Does This Ace Algorithm Solve
First, look at the human part: when the hand is blocked, can the movement still be tracked accurately?
When people work in the real environment, their upper bodies are constantly moving, and their hands will cross or be blocked by objects. The tracking system needs to restore the body and hand posture as accurately as possible from the constantly changing picture.
But in the video demonstration, even under such circumstances, the algorithm tracking can still maintain high accuracy.
Behind this, the self-developed MaxEgo undertakes the extraction of hand and posture information, which is an important part of this processing capability.
Tracking of occluded hands remains stable
The difficulty is that real operations will not always keep both hands in front of the lens for the convenience of the algorithm.
Fingers are blocked by objects during grasping, and the perspective changes when turning around. These clips contain valuable action information. The more reliable the tracking is, the more likely subsequent training will retain these details.
Then look at the space part: after the person walks out and comes back, can the system still align to the original starting point?
This tests SLAM, which means determining its own position while building a map of the surrounding environment.
When the photographer keeps moving, the positioning error may gradually accumulate. When the person clearly returns to the starting point, the trajectory recorded by the system may not necessarily match it.
In this demonstration, the trajectory can be well closed and re-aligned with the starting point after going out and walking around.
As a result, the hand movements have a more complete spatial background —
How the hand moves, where the person is, and the geometric relationship of the surrounding environment can be understood together.
This all relies on the solid underlying skills of the team.
The core engineering team of Maxinsights brings together senior talents from top companies such as Waymo, Google, Meta, QCraft, and cutting-edge laboratories of Stanford and Berkeley, most of whom have deep experience in autonomous driving for many years.
As an industry driven by massive data, the team already has mature practical experience in managing and iterating millions of levels of vehicle-end data.
Now turning to robotics, the underlying engineering methods and data links are still fully proficient.
Finally, look at the movement part: after the trajectory is tracked, we also need to clearly explain what the person is doing.
In the demonstration, the language labels are generated by the self-developed video understanding and labeling model, which can accurately correspond to the actual operations.
In Maxinsights' data system, this information is organized in layers according to environment, task, subtask and action instruction.
For example, an operation in a hotel room can be broken down into:
Hotel room → Clean the room → Place the kettle → Put the kettle on the table with both hands.
This connects the motion trajectory with the task meaning: the model can not only obtain the information of how the body and hands move, but also know what goals these actions are achieving.
With the three layers of information of how the body moves, how to move in space, and what is being done, a daily video has more complete learning value.
The deeper the data is processed, the richer the information available for the model to learn from the original picture.
According to the internal version evaluation, from MaxEgo 0.5 to 1.0, the key error indicators on related difficult cases have dropped by more than half.
What supports the algorithm iteration is more than 2 billion frames of human posture data, as well as a verification system composed of manually corrected and high-precision reference data.
This also explains why data accumulation and algorithm capabilities promote each other.
In addition to posture tracking, spatial positioning and motion understanding, there is deeper physical information to be mined.
According to Maxinsights' sorting, there are still body and environment geometry, object pose, and the most difficult to measure and most valuable contact information at higher levels.
Therefore, in Maxinsights' process, collection is followed by quality inspection, format unification, video understanding, spatial processing, posture tracking and manual correction, and then inspection and format conversion are carried out before delivery.
Two models plus four-stage automated QA let the machine complete 98% of the quality inspection coverage, and humans only deal with the parts that the machine cannot handle accurately.
The samples corrected by humans will flow back to become the training data for the next version of the model.
This is a flywheel that works better and better: the more deliveries you make, the more accurate the model is, and the cheaper the labeling is.
The Skills of 5000 People Are "Digitized" For the First Time
Beyond technology, what is more difficult for Maxinsights to replicate is another thing —
How to get more than 5000 collectors around the world to continuously shoot real life in different ways.
Maxinsights has deployed directly managed on-site teams and community networks in multiple countries, covering more than 2000 collection environments and over 50000 scenarios, with a monthly collection capacity of 450,000 hours.
The operation is scheduled by an Agent platform:
- It will monitor the data distribution, find that "there are enough restaurant kitchens, but not enough repair workshops", and then allocate equipment and manpower to the gaps;
- It will identify invalid clips, posed clips, and wearing offset;
- It will also rotate the scenarios every two weeks to prevent the same kitchen from being shot so many times that the model gets tired of it.
Who are the collectors?
Many of them are housewives, restaurant chefs, workshop workers, farmers who pick different fruits on farms, engineers who repair various cars, caregivers who take care of the elderly, etc......
The feel of a housewife who has been cutting, sorting vegetables and unscrewing bottle caps for decades used to only exist in her own muscle memory.
Now, when she wears a camera to cook a meal, these movements are decomposed into hand trajectories, action sequences and semantic labels, becoming the first lesson for robots to learn to do housework.
This is probably the first time that ordinary people can sell their "craft skills" directly to Physical AI.
For model companies, this is also the most irreplaceable part —
The mess, mistakes and on-site adjustments in real life can never be acted out in the laboratory no matter how hard you try.
Millions of Hours Enter Model Training
The two most popular robots this summer are "fed" with Maxinsights data.
On August 10, Dyna Robotics released Dyna-2, which used 1 million hours of human first-person video for pre-training.
This is equivalent to a person living with his eyes open for 170 years.
From 1000 hours to 1 million hours, the model performance continues to improve monotonically across four orders of magnitude, and no ceiling has been seen yet. Moreover, the model pre-trained with more human data also performs better on robot data that has never been seen before.
Maxinsights is the largest supplier of this 1 million hours of data.
Back in May, Genesis AI released GENE-26.5, where the robot uses both hands to complete a 20-step meal: cutting vegetables, beating eggs, pouring smoothies, without human intervention in the middle.
Maxinsights is also the first-person data partner of GENE-26.5.
Genesis superimposes first-person video and data from sensor-equipped gloves, and the result is that many high-difficulty skills can achieve autonomous execution with less than one hour of robot-specific data.
The two embodied intelligence models that became popular in this round in Silicon Valley are both