To understand Galaxy Universal is to fully grasp the future of embodied intelligence.
How should we define the ChatGPT moment of embodied intelligence?
On August 19, at the site of the World Robot Conference (WRC), during the sub-forum "The Evolution Road of Embodied Intelligence Large Models" hosted by Galaxy Universal, Wang He, Founder and CTO of Galaxy Universal, raised this question to the audience.
His answer is that robots can achieve zero-shot generalization for common skills that they have not specifically learned, and ordinary users can teach them new movements without an algorithm background. To achieve this, humanoid robots need a brain that can understand the physical world, control the whole body and hands, and learn continuously.
This vision has been given a new physical carrier on Galbot ET1 ("Galaxy Star Kid"), the latest small humanoid robot released by Galaxy Universal. According to the introduction, "Galaxy Star Kid" is the world's first intelligent humanoid robot with autonomous learning capability. It is equipped with AstraBrain-Agent independently developed by Galaxy Universal, which can observe the environment in real time, understand human instructions, and dynamically plan behaviors. Humans do not need to prepare action videos to teach it new motor skills through interaction.
Behind the flexible body of "Galaxy Star Kid" is the latest evolution of its embodied large model, AstraBrain. While the embodied intelligence industry is still in fierce competition around physical performance, Galaxy Universal has placed heavier bets on embodied large models, which are also determining the upper limit of robots.
This is also why Galaxy Universal has become the most concerned company at this WRC.
When Agent Enters the Physical World
The reason why "Galaxy Star Kid" has aroused widespread discussion is that it carries a new product logic.
Over the past two years, Agent has become one of the most popular concepts in the AI industry. Agents in the digital world can call tools between software, help users find information, and perform online tasks on behalf of people. Physical agents face the ever-changing real environment. No matter the change of object shape, or the change of light and position, the model needs to understand it quickly, and then drive the body to respond.
Galaxy Universal defines AstraBrain-Agent as a "physical world native agent". Its perception and planning are directly oriented to the real space, and the results of thinking will eventually be transformed into body movements. After the robot performs the action, the environment changes accordingly, and new feedback will affect the next judgment.
The interactive self-learning capability of "Galaxy Star Kid" is generated under this logic. It can recognize the motion trajectory of human dancers in real time and follow them to complete street dance movements. Facing difficult movements such as supporting the body on the ground and handstands, it can still maintain physical stability. This flexible physical performance comes from the support of AstraBrain-WBC, the general cerebellum of AstraBrain.
Behind this model, there are more than 100,000 hours of human motion data. Part of it comes from high-precision motion capture, and the other part is converted from human videos on the Internet. With the help of these data, the model can learn how humans maintain balance, coordinate joints and complete coherent movements, and then transfer the motion ability to the robot, without writing the whole set of motion trajectories in advance.
Behind this is the evolution of embodied large model technology. In the past, to add a new skill to a robot, a professional team usually needed to recollect data and complete training, and the skill was easily bound to a specific ontology and scenario. AstraBrain-Agent hopes to retain the general capabilities that the model has already acquired, and then adapt to new tasks with a small amount of interaction. Every time the robot learns a new skill, it is continuously expanding the capability of the same brain.
Around AstraBrain, Galaxy Universal has built a whole-brain architecture including the cerebrum, pons and cerebellum. The cerebrum is responsible for understanding the environment and planning actions, the cerebellum is responsible for controlling the whole body and hands, and the pons transmits high-level intentions to specific actions.
Among them, the cerebrum adopts the World-Action Model WAM architecture first proposed by Galaxy Universal. VLA can generate actions based on vision and language, and the world model is good at deducing how the environment will change. WAM integrates the two capabilities into a unified model, allowing the robot to understand the physical results that an action may bring, and then decide the next behavior accordingly.
The core breakthrough of AstraBrain WAM is to integrate cross-ontology, cross-scenario and multi-task capabilities into the same base model. The same brain can drive the G1 robot to use dexterous hands to pick up goods in supermarkets, and also complete flexible operations such as folding clothes; after replacing the robot ontology, it can also perform new box-moving tasks. The accumulation of models is no longer limited by hardware and scenarios, and this extensive capability migration is the first of its kind in the industry.
"Galaxy Star Kid" is also equipped with the AstraBrain-WBC general cerebellum base model independently developed by Galaxy Universal. Trained on large-scale human motion data, it can convert the intentions generated by AstraBrain-Agent into stable and coherent body movements. The model also has excellent generalization ability when facing movements that have not appeared in training.
Tennis is one of the scenarios that can best test this set of technical capabilities. This sport is highly confrontational, leaving extremely short reaction time for robots. It needs to predict the trajectory of the incoming ball, adjust the position in time, and hit the ball accurately while maintaining balance. Equipped with Galaxy Universal's LATENT high-dynamic tennis confrontation regulation framework, "Galaxy Star Kid" has been able to adjust its movements in real time according to the incoming ball from real people in actual demonstrations, completing human-like sparring.
From following humans to making independent judgments, the industry has gradually formed a new consensus: the upper limit that a humanoid robot can reach depends more and more on how many new capabilities its "brain" behind it can develop. The more bodies it enters, the richer the experience it accumulates in the physical world, and the next robot will have a higher starting point accordingly.
A Real Test of One Brain
At the booth of Galaxy Universal, this "AstraBrain" has also been put into different tasks for testing.
The most immersive one is a seemingly ordinary breakfast task: the robot needs to complete the steps of taking bread, pouring water and arranging plates continuously. Accidental situations often occur during the execution, such as the cup being taken away by the audience, or the target object being suddenly blocked. However, in the on-site demonstration, the robot will re-plan according to the new state in front of it, and complete the remaining work smoothly.
Traditional robots execute preset processes in a fixed environment. As long as the conditions of a certain step change, the whole task is easily interrupted, but the real world rarely strictly follows the pre-written script. A person reaching out at the table, or a cup being moved away temporarily, will change what the robot should do in the next second.
The difficulty of long-horizon tasks lies in these temporary changes. The longer the movement lasts, the easier the errors generated earlier will affect the subsequent steps. The robot needs to constantly update its understanding of the environment during execution, while maintaining the memory of the whole task.
Another difficult scenario is folding clothes, which greatly tests the understanding of embodied large models on flexible objects. Clothes do not have a fixed shape, and every time the robot grabs them, the folds of the fabric will change. The model needs to re-judge the current state of the clothes and find a suitable grabbing point. Whether the task can continue depends on whether the model can update its judgment after each operation.
Of course, for the real industrial scenarios, one successful demonstration is far from enough. The accuracy rate in a large number of repeated operations every day determines whether customers are willing to pay for robots continuously.
Galaxy Universal is the most qualified embodied manufacturer to answer this question. The company has applied its products on a large scale in smart pharmacies and instant retail. Robots need to accurately find targets from tens of thousands of products, and hand them over to riders after grabbing. Such tasks are repeated many times every day, and accuracy, operation stability and exception handling capabilities directly affect business operations.
In industrial scenarios, the dual arms of Galaxy Universal's Galbot S1 have a maximum load of 50 kg, which can undertake heavy material handling. After the load increases, the robot needs to readjust its body posture and force output mode, and also judge the position of surrounding personnel to ensure collaborative safety.
Robots used in different scenarios have different forms, and the problems they face vary greatly. On one side are soft and easily deformed clothes, on the other side are dense shelves, and on the other side are heavy industrial objects. They share the technical base of AstraBrain, which also verifies the generalization ability of AstraBrain. This is exactly the starting point for embodied large models to have general capabilities.
For Galaxy Universal, the deployment of models in the industry has another value. Smart pharmacies, instant retail and industrial production lines will continuously generate real data. These data record problems that are difficult to cover in the simulation environment, such as changes in packaging materials, deviations caused by long-term operation of equipment, and random interference brought by on-site personnel. After the model is trained in a new round, the updated capabilities will be returned to the robot.
At the forum, Galaxy Universal also announced that it will open the simulation platform, data acquisition equipment, embodied base model and reinforcement learning post-training pipeline to technology enterprises and industrial partners. Ordinary developers can also create new movements and applications around "Galaxy Star Kid". This is also the inevitable path for embodied large models to become general: the more open the ecosystem is, the more participants there are, the richer the real problems that the large model contacts, and the faster the evolution speed of the model will be.
Only when the model enters the real world can the technical achievements be transformed into productivity that can grow continuously. Galaxy Universal has proved at this WRC that AstraBrain already has such capabilities.
Embodied Large Model Is the New Watershed
In the past few years, the spotlight of the humanoid robot industry has mostly been on the body.
Unitree Robotics, which stands out in robot motion control and engineering capabilities, its debut on the stock market on August 19 is a case in point. On the opening day, Unitree Robotics achieved an increase of more than 629%, and its market value once reached 444.9 billion yuan. The capital market gave a high valuation to the commercial value of humanoid robot ontologies with a very warm debut.
Robots that can run fast, jump high and withstand huge impacts are indeed the most intuitive measure of technological progress. But when robots enter pharmacies, supermarkets and factories, a "movable" robot obviously can no longer meet the needs. It needs to understand instructions with vague meanings, handle objects that have never been seen in training, and cope with the movement of people and the displacement of items. The body determines where the robot can reach, and the brain determines its capability boundary.
Thus, embodied large models have become a new watershed in the industry.
The performance of the ontology can be improved faster through supply chain and engineering investment, but the growth cycle of a general brain is often longer. It needs to learn physical laws in a large number of tasks and be tested in real scenarios. The richer the tasks the model experiences, the more experience it can call when dealing with new problems. Even if latecomers adopt similar hardware, it is difficult to quickly make up for this kind of know-how.
The listing of Unitree demonstrates the commercial value of robot bodies, while the competition of embodied large models is defining the capability coordinates of the next stage of the industry.
Galaxy Universal regarded data infrastructure as the core of embodied large models very early. To this end, AstraData has built a five-layer data pyramid. Among them, Internet data helps the model understand semantics, human motion data provides operational experience, the simulation platform can generate training samples on a large scale, and the real robot teleoperation data is responsible for calibrating fine movements. After the robot enters the actual scenario, the problems generated in the work will be returned to the training system.
Wang He disclosed at the forum that Galaxy Universal has accumulated 1 million hours of human data and 80,000 hours of real scenario return data. As early as 2021, the team began to build a first-person human-object interaction dataset. Today, data acquisition equipment, simulation platforms and model evaluation systems have been connected to the same set of infrastructure.
When more robots enter the real scenario, AstraBrain can obtain richer physical experience; after the model capability is improved, the adaptation cost of new robots and new tasks will also decrease. The deployment scale and model capability promote each other, forming a continuously growing data flywheel.
Back to the most discussed issue at this WRC: When will the ChatGPT moment of embodied intelligence arrive?
The standard given by Wang He is that robots can achieve zero-shot generalization for common skills that humans do not need to learn specifically, with a success rate of 70% to 80%; ordinary users can also complete post-training at low cost, so that the model can quickly adapt to their own working environment. At this time, embodied intelligence will truly have a popularization foundation similar to ChatGPT.
Reaching 70% to 80% of the model capability is only the first step. In reality, customers still need algorithm engineers to collect robot motion data, then complete labeling and debugging. The high adaptation cost will keep robots out of the reach of a large number of small and medium-sized enterprises and ordinary users. To popularize embodied large models, it is necessary to solve the problem of "anyone can teach and anyone can use".
In his speech at this forum, Wang He also introduced the WAM-TTT test-time training scheme. Users wear a first-person camera to shoot the process of completing their work, and the model can deploy with unlabeled videos, without recollecting robot motion data.
This further promotes the generalization of embodied intelligence. The unified base model is responsible for accumulating general knowledge of the physical world. Humans only need to demonstrate their work, and robots can quickly obtain corresponding capabilities. When the process of teaching robots is simple enough, embodied large models can truly enter thousands of industries.
At the end of the forum speech, Wang He depicted a broader industrial picture: In the future, robots may have a shipment scale close to that of mobile phones, maintain the product value of the automobile level, and inherit the continuous upgrading capability of large models at the same time. After one set of model completes evolution, countless robots can get capability improvement. This kind of scale benefit brought by software will open up a trillion-level market.
Humans have spent decades accumulating language and images for the digital world, and the training corpus of the physical world still needs to be written by robots themselves. They will identify tens of thousands of products in pharmacies, handle accidents in factories, and learn a new job in the demonstration of ordinary people. Every real action will become the starting point of the next evolution.
When the same brain can enter different bodies, and when teaching a robot a skill becomes as natural as teaching a human, the ChatGPT moment of embodied intelligence will truly arrive. A brain that can share experience and evolve continuously in thousands of bodies will allow AI to truly enter the physical world, and open the real entrance for embodied AGI.