If you understand Galaxy Universal, you will grasp the future of embodied intelligence.
How on earth should we define the ChatGPT moment of embodied intelligence?
On August 19, at the site of the World Robot Conference (WRC), at the sub-forum "The Evolution Road of Embodied Intelligence Large Models" hosted by Galaxea, He Wang, Founder and CTO of Galaxea, raised this question to the audience.
His answer is that robots can achieve zero-shot generalization for common skills that they have not specifically learned; ordinary users can teach it new movements without an algorithm background. To achieve this, humanoid robots need a brain that can understand the physical world, control the whole body and both hands, and keep learning.
This vision has been given a new physical carrier on Galbot ET1 ("Galaxea StarKid"), the latest small humanoid robot released by Galaxea. It is introduced that "Galaxea StarKid" is the world's first intelligent humanoid robot with autonomous learning capability. It is equipped with AstraBrain-Agent independently developed by Galaxea, which can observe the environment in real time, understand human instructions, and dynamically plan behaviors. Humans can teach it to learn new movement skills through interaction without preparing movement videos in advance.
Behind the agile body of "Galaxea StarKid" is the latest evolution of its embodied large model — AstraBrain, the Galaxea StarBrain. While the embodied intelligence industry is still engaged in fierce competition around physical performance, Galaxea has placed heavier bets on embodied large models, and this model is also determining the upper limit of robots.
This is also the reason why Galaxea has become the most concerned company at this WRC.
When Agents Step into the Physical World
The reason why "Galaxea StarKid" has sparked widespread discussion is that it carries a new product logic.
Over the past two years, Agents have become one of the most popular concepts in the AI industry. Agents in the digital world can call tools between software, help users find information, and perform online tasks on behalf of people. Physical agents face a constantly changing real environment. Whether it is the change of object shape, or the change of light and position, the model needs to understand quickly and then drive the body to respond.
Galaxea defines AstraBrain-Agent as a "physical world native agent". Its perception and planning are directly oriented to the real space, and the results of thinking are eventually transformed into physical movements. After the robot performs the movement, the environment changes accordingly, and new feedback will affect the next judgment.
The interactive self-learning capability of "Galaxea StarKid" is generated under this logic. It can recognize the movement trajectory of human dancers in real time and follow them to complete street dance movements. Facing difficult movements such as supporting the body on the ground and handstands, it can still maintain physical stability. This agile physical performance comes from the support of AstraBrain-WBC, the general cerebellum of Galaxea StarBrain.
Behind this model, there are more than 100,000 hours of human movement data. Part of it comes from high-precision motion capture, and the other part is converted from human videos on the Internet. With the help of these data, the model can learn how humans maintain balance, coordinate joints and complete coherent movements, and then transfer the motion capability to the robot, without writing the whole set of movement trajectories in advance.
Behind this is the evolution of embodied large model technology. In the past, to add a skill to a robot, a professional team usually needed to collect data again and complete training, and the skill was easily bound to a specific ontology and scenario. AstraBrain-Agent hopes to retain the general capabilities that the model has already acquired, and then adapt to new tasks with a small amount of interaction. Every time the robot learns a new skill, it is continuously expanding the capability of the same brain.
Around Galaxea StarBrain, Galaxea has built a whole-brain architecture including the cerebrum, pons and cerebellum. The cerebrum is responsible for understanding the environment and planning actions, the cerebellum is responsible for controlling the whole body and both hands, and the pons transmits high-level intentions to specific movements.
Among them, the cerebrum adopts the World-Action Model WAM architecture first proposed by Galaxea. VLA can generate actions based on vision and language, and the world model is good at deducing how the environment will change. WAM integrates the two capabilities into a unified model, allowing the robot to understand the physical results that an action may bring, and then decide the next behavior accordingly.
The core breakthrough of AstraBrain WAM is to integrate cross-ontology, cross-scenario and multi-task capabilities into the same base model. The same brain can drive the G1 robot to use dexterous hands to pick up goods in supermarkets, and also complete flexible operations such as folding clothes; after replacing the robot ontology, it can also perform new box-moving tasks. The accumulation of the model is free from the restrictions of hardware and scenarios, and this extensive capability transfer is the first of its kind in the industry.
"Galaxea StarKid" is also equipped with the AstraBrain-WBC general cerebellum base model independently developed by Galaxea. Trained on large-scale human motion data, it can convert the intentions generated by AstraBrain-Agent into stable and coherent physical movements. The model also has excellent generalization capability when facing movements that have not appeared in training.
Tennis is one of the scenarios that can best test this set of technical capabilities. This sport is highly confrontational, leaving very little reaction time for the robot. It not only needs to predict the trajectory of the incoming ball, but also adjust its position in time, and hit the ball accurately while maintaining balance. Equipped with Galaxea's LATENT high-dynamic tennis confrontation regulation and control framework, "Galaxea StarKid" has been able to adjust movements in real time according to the incoming ball from real people in actual demonstrations to complete anthropomorphic sparring.
From following humans to making independent judgments, the industry has gradually formed a new consensus: the upper limit that a humanoid robot can reach depends more and more on how many new capabilities its "brain" behind it can develop. The more bodies it enters, the richer the experience it accumulates in the physical world, and the next robot will have a higher starting point accordingly.
A Real Test for One Single Brain
At Galaxea's booth, this "Galaxea StarBrain" was also put into different tasks for inspection.
The most immersive task is a seemingly ordinary breakfast task: the robot needs to complete a series of steps including taking bread, pouring water and arranging plates continuously. Accidental situations often occur during the execution, such as the cup being taken away by the audience, or the target object being suddenly blocked. But in the on-site demonstration, the robot will re-plan according to the new state in front of it and complete the remaining work smoothly.
Traditional robots execute preset processes in a fixed environment. As long as the conditions of a certain step change, the whole task is easily interrupted, but the real world rarely strictly follows the pre-written script. A person reaching out at the table, or a cup being moved away temporarily, will change what the robot should do in the next second.
The difficulty of long-horizon tasks is hidden in these temporary changes. The longer the movement lasts, the more likely the errors generated earlier will affect the subsequent steps. The robot needs to constantly update its understanding of the environment during execution, while retaining the memory of the whole task.
Another difficult scenario is folding clothes, which greatly tests the embodied large model's understanding of flexible objects. Clothes do not have a fixed shape, and every time the robot grabs them, the folds of the fabric will change. The model needs to re-judge the current state of the clothes and find a suitable grasping point. Whether the task can continue depends on whether the model can update its judgment after each operation.
Of course, for the real industrial scenarios, a single successful demonstration is far from enough. The accuracy rate in a large number of repeated operations every day determines whether customers are willing to pay for robots continuously.
Galaxea is the most qualified embodied manufacturer to answer this question. This company has applied its products on a large scale in smart pharmacies and instant retail. Robots need to accurately find targets among tens of thousands of products, complete grasping and then hand them over to riders. Such tasks are repeated a large number of times every day, and accuracy, operation stability and exception handling capabilities directly affect the business.
In industrial scenarios, the Galbot S1 dual arms of Galaxea have a maximum load of 50 kg, which can undertake heavy material handling. After the load increases, the robot needs to readjust its posture and force output mode, and also judge the position of surrounding personnel to ensure collaboration safety.
Robots used in different scenarios have different forms, and the problems they face vary greatly. On the one hand, there are soft and easily deformable clothes, on the other hand, there are dense shelves, and on the other hand, there are heavy industrial objects. They all share the technical base of Galaxea StarBrain, which also verifies the generalization capability of Galaxea StarBrain. And this is exactly the starting point for the embodied large model to be universal.
For Galaxea, the application of the model in the industry has another value. Smart pharmacies, instant retail and industrial production lines will continuously generate real data. These data record problems that are difficult to cover in simulation environments, such as changes in packaging materials, deviations caused by long-term operation of equipment, and random interference brought by on-site personnel. After the model goes through a new round of training, the updated capabilities will be returned to the robot.
At the forum, Galaxea also announced that it will open its simulation platform, data acquisition equipment, embodied base model and reinforcement learning post-training pipeline to technology enterprises and industrial partners. Ordinary developers can also create new movements and applications around "Galaxea StarKid". This is also the inevitable path for embodied large models to become universal: the more open the ecosystem becomes, the more participants there are, the richer the real problems the large model is exposed to, and the faster the model will evolve.
Only when the model enters the real world can technical achievements be transformed into productivity that can grow continuously. Galaxea has proved at this WRC that Galaxea StarBrain already has such capabilities.
Embodied Large Model Is the New Watershed
In the past few years, the spotlight of the humanoid robot industry has mostly been on the body.
Unitree Robotics, which stands out in robot motion control and engineering capabilities, its public debut on August 19 is a case in point. On the opening day, Unitree Robotics achieved an increase of more than 629%, with a market value of reaching 444.9 billion yuan at one point. The capital market gave a high price for the commercial value of humanoid robot ontology through a very warm debut.
Robots that can run fast, jump high and withstand huge impacts are indeed the most intuitive measure of technological progress. But when robots enter pharmacies, supermarkets and factories, a "mobile" robot obviously can no longer meet the demand. It needs to understand instructions with vague meanings, handle objects that have not been seen in training, and cope with the walking of personnel and the displacement of objects. The body determines where the robot can reach, and the brain determines its capability boundary.
Thus, embodied large models have become a new watershed in the industry.
Ontology performance can be improved faster through supply chain and engineering investment, but the growth cycle of a general brain is often longer. It needs to learn physical laws in a large number of tasks and enter real scenarios for inspection. The richer the tasks the model experiences, the more experience it can call on when dealing with new problems. Latecomers, even with similar hardware, will find it difficult to quickly make up for this kind of know-how.
The listing of Unitree demonstrates the commercial value of robot bodies, while the competition of embodied large models is defining the capability coordinate of the industry in the next stage.
Galaxea regarded data infrastructure as the core of embodied large models very early. To this end, Galaxea StarData has built a five-layer data pyramid. Among them, Internet data helps the model understand semantics, human movement data provides operational experience, the simulation platform can generate training samples on a large scale, and real robot teleoperation data is responsible for calibrating fine movements. After the robot enters the actual scenario, the problems generated in the work will return to the training system.
He Wang disclosed at the forum that Galaxea has accumulated 1 million hours of human data and 80,000 hours of real scenario return data. As early as 2021, the team began to build a first-person human-object interaction dataset. Today, data acquisition equipment, simulation platform and model evaluation system have been connected to the same set of infrastructure.
When more robots enter the real scenario, Galaxea StarBrain can obtain richer physical experience; after the model capability is improved, the adaptation cost of new robots and new tasks will also decrease. The deployment scale and model capability promote each other, forming a continuously growing data flywheel.
Back to the most discussed issue at this WRC: when will the ChatGPT moment of embodied intelligence arrive?
The standard given by He Wang is that robots can achieve zero-shot generalization for common skills that humans do not need to learn specifically, with a success rate of 70% to 80%; ordinary users can also complete post-training at low cost, so that the model can quickly adapt to their own working environment. At this time, embodied intelligence will truly have a popularization foundation similar to ChatGPT.
Reaching 70% to 80% of the model capability is only the first step. Customers in reality still need algorithm engineers to collect robot movement data, then complete labeling and debugging. The high adaptation cost will keep robots out of the reach of a large number of small and medium-sized enterprises and ordinary users. To popularize embodied large models, it is necessary to solve the problem that "everyone can teach and everyone can use".
In his speech at this forum, He Wang also introduced the WAM-TTT test-time training scheme. Users wear a first-person camera to shoot the process of themselves completing the work, and the model can use unlabeled videos for deployment without re-collecting robot movement data.
This has taken another step forward in the generalization of embodied intelligence. The unified base model is responsible for accumulating general knowledge of the physical world. Humans only need to show their own work, and robots can quickly obtain the corresponding capabilities. When the process of teaching robots is simple enough, embodied large models can truly enter thousands of industries.
At the end of the forum speech, He Wang depicted a far-reaching industrial picture: In the future, robots may have a shipment scale close to that of mobile phones, maintain product value at the automotive level, and inherit the continuous upgrading capability of large models at the same time. After one set of model completes evolution, countless robots can get capability improvement. This kind of scale benefit brought by software will open up a trillion-level market.
Humans have spent decades accumulating language and images for the digital world, and the training corpus of the physical world still needs to be written by robots themselves. They will identify tens of thousands of products in pharmacies, handle accidents in factories, and learn a new job from the demonstration of ordinary people. Every real action will become the starting point for the next evolution.
When the same brain can enter different bodies, and when teaching a robot a skill becomes as natural as teaching a human, the ChatGPT moment of embodied intelligence will truly arrive. A brain that can share experience and evolve continuously across thousands of bodies will allow AI to truly enter the physical world and open the real entrance for embodied