HomeArticle

Kevin Kelly × Wang He: Humanoid Robots Are Stepping into the Era of Cerebrum and Cerebellum Coordination

未来可栖2026-09-28 16:28
Opportunities, Limitations and Security Boundaries

"If we look at it from the perspective of model development, China and the United States can be said to have their own strengths, and are moving towards each other. China has a more complete industrial ecosystem, with a faster iterative closed loop covering models and data; the United States takes the lead in pure large model technology, and it can even endow a text model with embodied capabilities."

On September 21, at the salon "Meet at the Peak - Kevin Kelly Talks with New Forces of Chinese Tech" co-hosted by 36Kr and Chen Yuan 139 of Zhongjian Zhidi, Wang He, Founder and CTO of Galaxy Universal, put forward the above views.

Event site

Talking about specific technological breakthroughs and concepts, Kevin Kelly also affirmed them: "This is exactly your mission." KK also said that he can imagine China becoming a leader in AI-related fields including embodied robots.

The "AstraTennis Moment" of Embodied Intelligence

The first time AI shocked the world was in 2016. That year, AlphaGo defeated the top human player in the Go world, making the world truly feel the power of artificial intelligence for the first time. This also brought great shock to Wang He, who was a student at Tsinghua University at that time. But he also admitted that at that time, he still did not imagine that the GPT paradigm would emerge only three years later; and by 2022, ChatGPT could rapidly extend the intelligence specialized in Go to general intelligence that masters any language.

At the opening ceremony of the World Humanoid Robot Games on August 22, Galaxy Universal launched the world's first end-to-end model for humanoid robots oriented to fully autonomous tennis confrontation - AstraTennis. Through a wonderful human-machine tennis duel, it proved that embodied intelligence can complete perception, game and whole-body movement execution in the physical real space full of uncertainties.

In this regard, US investor Andrew Kang once commented: "AlphaGo for every sport is coming."

Wang He confessed that when the team proposed to let humanoid robots play tennis fully autonomously in June 2025, he himself was skeptical - tennis has extremely high requirements for hitting accuracy, landing point, movement and balance. Until the eve of the Spring Festival, the team completed mature training in the simulation world, and he decided to rent a real venue for a week to verify. "Finally, the research progress refreshed my cognition again: it can play tennis."

Wang He

He also mentioned that in addition to whole-body movement, Galaxy Universal has also made a breakthrough in the dexterous hand operation of humanoid robots, which is also the first in the world. "Behind the continuous challenges, it is the optimism mentioned by Mr. KK that supports us."

Talking about the development of humanoid robots and embodied intelligence, Kevin Kelly also said, "Manufacturing humanoid robots will be the most complex transformation we have ever done. It is far more complex than aviation and factories. Coupled with the complexity of chips, the complexity and sensitivity of hands, as well as the power and computing power scheduling required to run it, this is the most complex thing we have built. That's why it is the most difficult machine to make well."

On a global scale, the "breakthrough" of embodied intelligence is still ongoing.

This September, OpenAI released its flagship model GPT-6 Astra, which is officially called "the most intelligent and most aligned model to date".

The Galaxy Universal team found through evaluation that GPT-6 Astra shows strong reasoning ability in embodied tasks. In the RoboDojo experiment they conducted, GPT-6 Astra+π0.5 achieved an average score of 62.6 in ten tasks, nearly double that of the second place. In another RoboLab experiment, the success rate of GPT-6 Astra Direct in 50 experiments was 98%; the success rate of Astra+π0.5 was 92%. While π0.5, DreamZero and others in the same group of statistics were all less than 40%.

"Astra is the first model in the GPT series that demonstrates embodied capabilities. Its embodied ability is very strong, and it is very 'smart' - no special training is needed, you can directly command it with language, and it can execute. For example, randomly ask it to build a tower with wooden blocks and planks. You can see that GPT-6 Astra will keep thinking about what to do at each step, and decide whether to call the embodied model π0.5 or complete it by itself. It can do this completely untrained task for the first time. Such examples are numerous, and the success rate of all tasks can exceed 60%, while the best performance of previous models is only about 10%, which is exactly the breakthrough significance of it." Wang He commented on it.

What makes GPT-6 Astra truly powerful is not only that it can "understand", but also that it can find "what is wrong". For example, in the experiment, after it finds that the grasping posture is not ideal, it will change the approaching method; during the packing process, it can find that there are missing objects behind the box...... This shows that GPT-6 Astra does not mechanically execute a pre-generated trajectory, but forms a closed loop of "observation - decision - execution - re-observation".

Enable Robots to Realize the Coordination of Cerebrum and Cerebellum

"Digital intelligence is a kind of knowledge smart, while robot intelligence is a kind of intuitive intelligence. The vast majority of human reasoning and thinking belongs to the slow system in the brain. Human operation belongs to the fast system in our brain, which is System 1. Such a fast system requires a lot of practice to turn such abilities into an innate, unconscious reaction. Large models can use a very complex thing that we call Chain of Thought, which is a completely different kind of intelligence.

For the next direction, Wang He said, "On this road, we also hope to build step by step a bionic brain system that can connect the cerebrum, cerebellum and pons like humans, which can not only make high-level judgments, but also realize low-level humanoid and bionic control, so that our humanoid robots can truly have intelligence."

In terms of technical routes, Wang He introduced two paths for large embodied models: Google's VLA (Vision Language Action) emphasizes that each data must have robot action output; OpenAI's World Model is more general, but the data does not necessarily contain robot-related content. Galaxy Universal took the lead in integrating the two and proposed the World Action Model (WAM) - which can predict both the future and actions. This technical route has now been widely recognized by the academic community.

With the support of the general brain WAM, Galaxy Universal's robot also completed a 30-minute full family housework challenge at the Humanoid Robot Games, covering the whole process of item placement, garbage cleaning, laundry drying and folding, and it can also respond to unexpected instructions such as picking up express delivery halfway, and finally won the gold medal of the event with almost full marks.

Wang He

In addition to the brain model, the cerebellum model of Galaxy Universal is also noteworthy. Wang He said: "In the past, when robots danced, we needed to import human movement trajectories into the robot in advance, adjust them and then deploy. Now our general cerebellum model can make the robot immediately repeat the dance after a human dances once. This post-training intelligence even surpasses human's ability to master limb movements, and it is really a general cerebellum model trained with 2 billion frames of human movement data."

For these breakthroughs, Kevin Kelly said, "The next major breakthrough will be to really get robots to work. What we have been doing will have a huge impact. So this is exactly where the excitement lies."

Facing the future, Wang He said: "We will turn our robot body, cerebrum and cerebellum, data platform, as well as the complete tool chain for training and deployment tools into four-in-one infrastructure, and open it to the whole industry. Everyone can use a very simple data collection process and post-training process to teach robots. At that time, humanoid robots, as a form of productivity and service capacity, will enter thousands of industries."

The Essence and Safety Boundary of Robots - They Are Partners, Not Substitutes

With the development of agents, in the future physical world, more and more robots will enter society, families and other human living scenarios. How should humans grasp the safety boundary between humans and robots? Will the much-discussed "AI killer" become a reality?

For the relationship between the two, Kevin Kelly believes that "when we give AI a body - that is, a robot, we need to understand: they are different from us, they are not human. The fact that they are different from human is not a defect, but precisely their biggest advantage. In my opinion, such intelligent carriers with physical entities are a valuable resource with low cost. Because there are many intellectual problems that cannot be solved by humans alone, nor by AI alone, but can be solved by the joint efforts of AI and humans. Humans and robots are such a partnership."

Event site

Wang He agrees with Kevin Kelly's point of view: The essence of a robot is AI, it is not a human. So I am quite optimistic about the future of robots. It will become human's partner, not a substitute or opponent of human."

Then what about the safety issues when robots enter human daily life in the future? Wang He took the already deployed pharmacy robot as an example: in deterministic closed scenarios such as factories and pharmacies, the rights and responsibilities are relatively clear. Galaxy Universal has deployed nearly a hundred robot pharmacy points across China, undertaking the work of dispensing medicine at night. To avoid the risk of wrong medicine dispensing, the team does not fully rely on the general reasoning of large models, but adds an additional deterministic verification program: the medicine must pass three verifications of text, image and barcode at the same time before it is allowed to be packed and out of the warehouse. In the past year and a half, it has processed nearly a million boxes of medicine in total, achieving zero wrong dispensing. In such scenarios, if losses are caused by robot failures, the responsible subject is the manufacturer.

Kevin Kelly believes that in the short term, AI companies must assume corresponding responsibilities, and supporting insurance mechanisms can also be established. Conversely, the AI companies that can survive are definitely those that can properly handle responsibility issues.

At the same time, Wang He also pointed out that in the completely open and complex family scenarios, there are full of unpredictable interference factors, and the risk complexity multiplies. "For example, when no one is at home, the robot knocks over the water, and the water droplets enter the power strip, causing a short circuit and fire, and it cannot clean up by itself. It is more dangerous than cats and dogs, because it can plug and unplug plugs and move all things in your house. If such a situation occurs in the future, I fully agree with what Mr. KK said, and I believe that by then we may use a set of systems to ensure safety."