GPT has dismantled the "embodied brain", will Embodied PhyStack be the next stop for embodied intelligence?
One of the most compelling narratives in embodied intelligence over the past two years has revolved around the "brain".
Overseas, from Physical Intelligence's π series, to Figure's Helix, and Google DeepMind's Gemini Robotics, a batch of globally high-profile embodied brain models are all working to integrate vision, language and motion into one unified learning framework.
Domestic startups are also flocking to this track. According to data from IT Juzi, in the first half of 2026, the domestic embodied intelligence sector recorded 322 financing deals totaling 935 billion yuan, a 5-fold increase compared with the same period in 2025. More than half of the capital flowed to VLA models, world models and data infrastructure, making the "brain-cerebellum-body" structure the mainstream technical path.
Yet after two years of exploration along this path, the industry has found that the "brain" remains an extremely tough nut to crack. However, the release of GPT-6 Astra on September 3, 2026 has brought new variables to the development of embodied intelligence.
OpenAI still officially positions it as a general-purpose model, with no direct mention of robots. But when researchers connected GPT-6 Astra to a humanoid robot, it instantly electrified the entire embodied intelligence industry. In the evaluation report GPT-6 Astra as an Embodied Policy published on GitHub on September 12, GPT-6 Astra outperformed all competitors with a crushing score and took the top spot on the RoboDojo ranking by surprise.
In GPT-6 Astra as an Embodied Policy, GPT-6 Astra ranks first by a large margin across ten robot tests
Are general large models really about to return to the "main table" of embodied intelligence?
1. General Large Models Return to the "Main Table"
To truly grasp the magnitude of the impact GPT-6 Astra has brought to the entire embodied intelligence industry, we first need to address two fundamental questions: What exactly is embodied intelligence trying to solve? And why is it so difficult?
The ultimate goal of embodied intelligence has always been unmanned operation, that is, to enable robots to achieve "autonomy". Making robots move is only the starting point of this journey.
Over the past decade, the fastest-progressing segments of the entire robotics industry have been the "body" and "control" layers. Robotic arms have solved repetitive operations in deterministic environments, and wheeled bases have expanded the operating radius from individual workstations to entire workshops... The diversification of body forms has allowed robots to reach more scenarios; the improvement of control capabilities has enabled robots to perform tasks stably in these scenarios. The combination of the two has gradually expanded the autonomy boundary of robots.
But the autonomy boundary of wheeled bases and robotic arms has an upper limit — they rely on relatively structured environments. Among all body forms, bipedal humanoid robots have greater potential to approach a more complete and extensive autonomy boundary. Because the human world is built around the bipedal form: door handles, stairs, tools and workstations are all designed for humans. After the industry realized this, humanoid robots have undergone continuous iterations on their body hardware for many years.
The control layer has also evolved in tandem. In the era of embodied intelligence, reinforcement learning and simulation training have promoted the motion control of humanoid robots. Walking, running, climbing stairs, and getting up after falling have gradually become basic capabilities of the industry. Dexterous hands, motors, reducers and force sensors are still improving, and the availability of robot bodies is also continuously rising.
The body and control layers are gradually maturing. But the real pain point has also emerged along with this progress: cognition and decision-making, that is, how to enable robots to independently judge what to do and how to do it in open environments, or in other words, to give robots a "brain".
PhyMotion, a Tsinghua-affiliated domestic startup, offers a typical case with its bipedal humanoid robot PHYBOT C2.
At the end of 2025, the PhyMotion team became the first in the world to enable a bipedal humanoid robot to play badminton autonomously against human opponents. At the Canton Fair in April this year, this 1.35-meter-tall bipedal humanoid robot held a racket and played intelligent badminton matches with on-site visitors, catching forehand hits, backhand hits and lift shots without any engineer intervention throughout the process.
(PhyMotion's bipedal humanoid robot plays badminton autonomously against a human opponent)
The difficulty of enabling a bipedal humanoid robot to play badminton autonomously against humans lies in the extremely narrow time window: the time from when the ball leaves the opponent's racket to when it arrives at the robot's hitting point is usually less than one second. The robot needs to complete the full workflow of observation, judgment, stepping, racket drawing and hitting within this period, and any delay in any link will lead to failure. The success of PHYBOT C2 proves that in a closed loop of tens of milliseconds, embodied intelligence can already rely on an end-to-end perceptual decision-making model to get rid of remote control and preset programs, and complete highly dynamic real-time interaction in autonomous state.
But it should also be noted that badminton is a task with relatively single goal and closed rules, where the robot does not need to think about what to do after the game. PHYBOT C2 has broken through the real-time decision-making bottleneck for humanoid robots, but the long-horizon task and multi-task planning capabilities of embodied intelligence have not yet been achieved. What is required for the latter is the ability to model the physical world, as well as the ability to understand human intentions and logical sequences.
In other words, it is truly very difficult to give robots a "brain".
In the early days of the large language model (LLM) boom, researchers did try to port large language models directly to embodied intelligent robots, but the results were far from satisfactory. In the Long Act benchmark test, which specifically evaluates planning-level autonomy in long-horizon household tasks, GPT-5 under the HoloMind framework only achieved a 59% target completion rate and a 16% full task success rate; Qwen3-VL-32B scored 51.2% and 15% respectively, while the success rate of the human control group in the same test was 93%.
As LLM paths failed to deliver expected results, paths like VLA and world models gradually became the industry's darlings. This was followed by large-scale data collection, larger VLA models, larger world models, and more powerful model architectures.
It is worth noting that VLA did not appear later than LLM. In 2023, Google RT-2 encoded robot motions into text tokens, allowing models to generate motions just like generating text, laying the foundation for the VLA path. Since then, driven by massive industry financing, "embodied brain" has become the hottest topic. Some industry insiders even joked that "you can't even call yourself a robotics company if you don't have a foundational model".
However, after massive capital injection, the problems have not disappeared. In the RoboDojo evaluation in July this year, the average success rate of 30 mainstream robot policies was only 12.8%, and the official comment stated that "most of them are still at the foot of the mountain getting used to the thin air". The "camera-foundation coupling" problem of VLA has not been solved: a slight change in camera perspective will cause the task success rate to drop sharply. Long-horizon task and multi-task planning capabilities are even more difficult to achieve.
Before that, the common solution in the industry was to collect data, collect more data, and use the Scaling Law to force out the "intelligence emergence" moment of the embodied intelligence industry.
That was until large models began to show new possibilities.
On September 3, 2026, OpenAI officially released GPT-6 Astra, which is officially advertised as a model "focused on complex reasoning, software engineering, computer usage, scientific research and professional documents". It has a context window of 1.05 million tokens and a maximum output of 128,000 tokens. The entire release document does not mention embodied intelligence at all.
In the official RoboDojo evaluation, GPT-6 Astra achieved an average success rate of 22.48% and a score of 28.97 across 42 robot tasks and 2100 tests, ranking higher than all 40 publicly available robot policies at that time, and outperforming the previous top-ranked DM0.5 which only scored 19.34% and 24.90.
2. The Entire Embodied Intelligence Industry is Undergoing Fundamental Changes
An architectural assumption that was previously widely accepted as the "endgame" has begun to falter.
In the past few years, the most common division in the industry was the "brain - cerebellum - body" structure: the upper layer is responsible for cognition and decision-making, the middle layer for motion control, and the hardware body for execution. With the rise of VLA and world models, the industry has invested massive resources trying to let the "brain" handle everything: it should understand the world, break down tasks, and directly command the robot body.
But GPT-6 Astra has demonstrated a new possibility: general large models themselves can take on the hardest part of the traditional "brain" — deep thinking.
The PhyMotion team believes that robotics companies no longer need to build a centralized all-powerful "brain" on their own. Deep thinking can be delegated to general large models, and what robotics companies need to supplement is the shallow thinking and real-time decision-making layer. The future embodied intelligence technology stack may evolve from three layers to four layers, which from top to bottom are general large model, physical agent PhyAgents, motion control system PhyCore, and robot body PHYBOT. PhyMotion names this architecture: Embodied PhyStack.
Schematic diagram of PhyMotion's Embodied PhyStack
The first layer, general large model: takes visual input, handles long-horizon tasks and multi-task planning, understands human intentions through deep thinking, breaks down long-horizon and multi-task sequences, judges logical order, and then invokes the capabilities of the physical agent PhyAgents to solve tasks. At the same time, the problem of slow response of cloud-based large models is compensated by PhyAgents.
The second layer, physical agent: takes visual input, performs real-time decision-making, with end-to-end VLA capability as the core. Through shallow thinking, PhyAgents translates the subtasks assigned by the large model into instructions that robots can execute in real time.
The third layer, motion control model: perceives the body state and performs real-time control, which is equivalent to the "cerebellum" of humans. PhyCore balances the body and drives the robot to execute according to the task instructions from PhyAgents.
The fourth layer, robot body: the PHYBOT hardware is responsible for execution, and based on changes in the real environment and body state, feeds back information to the above three layers for correction, re-planning and re-execution: gait and movement balance control issues are fed back to PhyCore for adjustment; environmental changes are fed back to PhyAgents for millisecond-level real-time decision-making; the overall execution rhythm and the reaching of capability boundaries are sent back to the first layer — the signals generated when the body touches the boundary allow the general large model to perform re-planning accordingly.
The significance of the Embodied PhyStack technology stack is far more than just adding an extra layer. It splits the "embodied brain" that used to be the full responsibility of a single company into a collaborative system between general large model providers and robotics companies: general large models are responsible for the deep thinking and reasoning planning they are best at, and robotics companies are responsible for the real-time decision-making, physical closed-loop and stable execution they are best at. For the embodied intelligence industry, this is a more realistic and cost-effective path: there is no need to build a super "brain" that does everything by itself, and let professionals do what they are best at.
PhyMotion, the company that proposed this path, has verified its feasibility through three sets of experiments.
3. PhyMotion's Three Sets of Experiments
The first set of experiments uses a general large model as the brain, paired with PhyMotion's PhyCore motion control system, to drive robots to perform long-horizon tasks.
In this set of experiments, the general large model acts as the deep thinking layer, directly pushing end-to-end VLA to long-horizon tasks. Without fine-tuning the model for specific tasks, it already achieves a sufficiently high completion rate. In the parts storage task, the "general large model + PhyCore" combination can already run through the full task workflow. Both in simulation environments and real robot tests, the robot can understand human intentions, combine the state of the real environment, and execute and complete long-horizon tasks and multi-tasks.
For example, in the battery installation simulation environment, the robot directly feeds back its failed actions to the general large model, uses the model's understanding of human intentions to think actively, finds the installation position, and completes the training task.
In the simulation environment, the robot uses the general large model to complete battery installation training
In real robot task execution, the robot will give feedback on failed actions based on its deep understanding of semantics and environment, and the general large model will re-think and re-plan to complete the task. In the experiment below, PHYBOT made two failed attempts when grasping a silver workpiece. Based on the general large model's understanding of deep intentions and the environment, it successfully completed the task on the third try.
The upper part shows the reasoning process of the general large model, the lower part shows the real robot completing the workpiece grasping process, at 200x speed
The first set of experiments verified the feasibility of large models in long-horizon tasks, but also exposed a problem: general large models are sufficient for deep thinking, but too slow for millisecond-level real-time decision-making. In