Unitree and Agibot share a unified brain, the mysterious model demo stuns the audience, featuring a 10-minute uncut one-shot footage.
Today's embodied intelligence companies produce demo clips that are as polished and exaggerated as commercials, but have you ever seen a completely uncut real video?
A few days ago, an industry insider friend sent me a 10-minute demo, shot in one continuous take with no editing, no off-site remote control, no manual instructions. It's even a bit rough, but the content presented made me slump down in my seat, and I couldn't calm down for a long time (doge).
In the video, the robot stretches its arm out of the window to clean the window, knows to move a box and step on it when it can't reach a high place, and can accurately resume a long-distance task even if it is interrupted midway...
What's more shocking is that the robot bodies of Unitree and Agibot appear in the video at the same time. These two "competitive products" with completely different hardware architectures, motion degrees of freedom and sensing systems share the same "brain", cooperate and collaborate closely with each other, making it a globally rare sample of a general-purpose brain across different robot bodies.
It is no exaggeration to say that this demo is enough to reshape the global embodied intelligence industry's perception and subvert the technical judgment of the track, and it may even rewrite the Scaling Law...
Where on earth did this "magic" technology come from?
Frame-by-frame breakdown of the video, all are iconic scenes
This video carries so much information. I carefully broke it down frame by frame, and the more I watched, the more fascinated I became.
Operating in enclosed spaces, with incredibly high stability
The shooting site is a rental house of about 15m², with dense furniture and narrow aisles that make it difficult even to turn around. In such an environment, there is basically no possibility of teleoperation (there is no space for operators to stand) or pre-set scripts.
From the very design of the environment, this video clearly tells everyone that it can only be a footage of robots that complete full autonomous learning, autonomous evolution and autonomous decision-making.
Two robots do housework in parallel in the same space, perceiving boundaries and planning paths in real time.
There is no collision, no getting lost, no stagnation throughout the whole process, perfectly adapting to the unfamiliar and complex real scene environment. Even the "battle-damaged" robot tied with a traction rope shows extremely strong flexibility.
The whole video has no cuts, no reshoots, and no manual instruction intervention.
Stretching out the window to clean the glass: the peak manifestation of dynamic capability and autonomous reasoning
From the very first task, the robot shows a "delicate" side that exceeds human imagination.
A Unitree robot uses a squeegee to clean the glass. Anyone who has used a squeegee knows that if you apply too little force, you can't clean it, and if you apply too much force, there will be abnormal noise. Cleaning glass with a squeegee tests force control and spatial perception more than cleaning with a rag, and the robot directly raises the difficulty to the hell level from the start.
After wiping a few times, the robot notices that one spot is still not clean, and it even judges on its own: the dirt may be on the outside of the window.
Then it performs a set of movements I have never seen before: turning sideways, leaning back, poking its head out, stretching its arm, and extending the hand holding the squeegee out of the window. (Stable as an old dog)
For some existing embodied models, achieving good force control is already their limit. But this robot, on the basis of precise torque control, seamlessly integrates the perception of the spatial environment (it does not hit the window frame when stretching out its hand) and dynamic calculation.
Not only does its single capability reach the industry ceiling, the model also demonstrates a unified three-in-one performance.
What makes me even more unable to calm down for a long time is that when it cannot clean the glass thoroughly, it immediately reasons and makes decisions to stretch its hand out of the window to clean it. This is the first time I have seen a model complete self-reasoning and self-evolution just like a human being.
This fully reflects that the model integrates vision, touch, dynamics and other capabilities into a unified perception, and has the ability of self-evolution. This time, I finally believe that robots can really do physical work, and what's more, they will do it more perfectly than humans.
Long-sequence task closed loop, no error accumulation allowed
After cleaning the glass, the Unitree robot smoothly puts the squeegee back in place, with accurate landing point and coherent movements.
Don't underestimate this finishing movement, which means it has completed the full-process closed loop of "fetching the tool - performing the task - storing the tool".
The Agibot robot next to it walks up to the washing machine, takes out the washed cushion and puts it on the sofa, then goes back to take out the washed doll, and accurately puts it into the wardrobe on the second floor, making sure every item is back where it should be.
The Agibot robot accidentally hits the glass, which further proves that this is the robot's autonomous decision-making, because the camera sometimes cannot handle the reflection of glass, but human teleoperation can (doge).
These two scenes are the epitome of the most basic capabilities of the entire 10-minute video. In long-sequence tasks, every task is completed smoothly, accurately and in one go, without accumulating errors or performing worse as the tasks become more and longer.
This solves the drawback that traditional VLA models are most afraid of long-sequence tasks.
Judging only from the movement performance, the model must adopt a brand new architecture. In the 10-minute ultra-long task chain, it continuously corrects errors and outputs stably, every movement is accurate and controllable, and its robustness is top-tier in the industry.
Tasks can be interrupted at will, and the model can resume execution from the breakpoint
Next, the ring of an alarm clock appears in the frame. (Turn up the volume and listen carefully, it's not very obvious)
At first I didn't understand the intention of the alarm clock. After watching it repeatedly, I found that the alarm clock is a pre-set reminder that will interrupt the two robots' current tasks, and make them start to perform special tasks of organizing the desktop and the refrigerator.
What's amazing is that after organizing the desktop and the refrigerator, the robots resume the tasks they were performing before the interruption, which shows that the model has the capabilities of arbitrary interruption, breakpoint resumption and task recovery.
Compared with the demos on the market that can only complete one task at a time and need careful protection, this rough video shows that in the real working environment, the robots are not only durable, but even have the super capability of "doing two things at once".
Then comes a complete long process of home storage: the Unitree robot takes the storage bag and puts it on the desktop, the Agibot robot sees the food on the refrigerator, judges that it should be stored in the refrigerator, immediately executes the task, and by the way takes out the food that should be thawed (I guess this mysterious team may be implying that robots will soon be able to cook four dishes and one soup).
There are no step-by-step instructions throughout the process, the robots autonomously break down tasks, plan movements, and complete the whole process from fetching objects to placing them step by step.
Nowadays, the vast majority of robots can only execute single-task serial processes, and when encountering interference midway, the task will basically collapse. However, this set of models has human-level task priority judgment and dynamic scheduling capabilities, and can switch task flows at any time. The intelligence demonstrated by such flexibility and stability far exceeds the capabilities of any known model on the market.
Agibot puts a scarf around Unitree's neck: the iconic scene of cross-body collaboration
Then, the climax of the whole video appears.
The two robots seem to have made an agreement that they must sort out the sundries on the desktop at one time, without moving back and forth.
So Unitree keeps placing objects on itself until its hands are full and it doesn't know how to operate. At this time, the Agibot robot walks over gently, slowly picks up a scarf on the table, and hangs it around the Unitree robot's neck.
The Agibot robot finds that the scarf is very long, so it holds the scarf with both hands. At this moment, Unitree seems to understand Agibot's intention, bends down, and Agibot wraps the scarf around Unitree's head and hangs it around the other's neck.
Come on dude! This smooth movement and the implicit interaction between their "eyes" even give off a sense of CP!
This video is not a pre-set division of labor, but a collaboration scheme independently explored by the model between the two robots of different bodies. Two robots of different brands independently judge their respective capability boundaries and complement each other perfectly, which is infinitely close to human collaboration, and even seems more tacit without language.
This is probably the first cross-body interaction in the world, and it may also be the first model in the world that can adapt to general robot bodies.
I guess this mysterious team wants to use two competitive robot bodies to tell everyone that this is the real general-purpose brain.
Hanging towels and carrying slippers: Are robots not only able to work, but have even started to "slack off"?
After Unitree, carrying a full load of items on its body, walks slowly into the distance, Agibot continues to store the remaining items.
At this time, it picks up a towel and tries to hang it on itself. After one, two, three attempts, it finally hangs the towel on its shoulder.
Three consecutive shocks!
The robot seems to know how to save the most effort: hanging the towel on the shoulder is more labor-saving than holding it in the hand.
The model can learn autonomously, and keep optimizing until it succeeds after failing once or twice.
The model's ability to understand its own body may be the core reason why it can be applied across different robot bodies.
This is not an isolated case.
Look carefully: when the robot sorts the slippers, it holds the hanging rope of the new slippers instead of the slipper body itself, because holding the rope is more labor-saving.
Another shock! While other robots are still being carefully polished just to complete work, this model not only gets the work done, but even directly learns to "slack off"? It's so smart.
Another climax: the robot learns to use tools, trying to use a box to raise its height
It's not over yet. The most shocking scene for me appears. In the process of putting all the items hanging on its body back in place, the Unitree G1 robot, which is only 1.3 meters tall, is stumped by the task of putting the scarf on the third layer of the cabinet.
A "divine" scene appears! It doesn't get stuck, nor does it give up. Instead, just like human reasoning, it finds a box by itself! Yes, it finds a box! It tries to stand on the box to raise its height and have another try.
It finds the box next to it, pushes it to the ground, tries to bend down to carry the box, but finds that it can't bend down.
After several attempts, just when I thought it was about to give up, it astonishingly kicks the box to the side of the cabinet with one foot!
The whole process reflects:
The capability of self-reasoning and decision-making. The robot knows how to reach the high objects, and how to move the box when it can't bend down.
The model keeps exploring its own physical capabilities. In the process of bending down and trying again and again, the robot seems to understand its own body boundary more and more, and makes choices that match its physical capabilities.
The process of self-evolution. Without any teaching, the robot even learns to use its foot to kick the box, and is good at using its own body to complete self-evolution in the environment.
This scene really makes my hair stand on end.
You know, the ability to use tools autonomously is a landmark feature of human beings.
But the robot in the video has really learned to use tools. It knows that the box can bear weight, that standing on it can increase its height, and even that it can use its foot to kick the box when it can't bend down.
Behind this set of movements is the integration of three capabilities: in-depth understanding of the rules of the physical world, continuous exploration of self-capabilities, and self-decision-making evolution.
Cross-body capability complementation, collaborative operation: from individual intelligence to group intelligence
Unitree kicks the box over but it goes awry. Just as it is thinking about how to correct the box, the 1.7m-tall Agibot Expedition A3 walks over. It seems to see Unitree's persistence and dilemma, puts down the small trolley in its hand (switches the task actively), and chooses to help its partner.
Just like the graduation medal-awarding ceremony, Unitree bends down and lowers its head, Agibot slowly takes the scarf off and puts it on the shelf.
Every step in the details reflects capabilities that outperform other models.
The two robots' understanding and cooperation with their own and each other's hardware capabilities go beyond individual intelligence and move towards group intelligence.
Agibot wraps the scarf around the back of Unitree's head before taking it off easily (you can imagine how other robots would yank this scarf off), showing the extreme understanding of force control and spatial perception.
The robot folds the scarf three times, which is an autonomous judgment on how to put the scarf into the wardrobe better. This is the real embodied "intelligence".
Finally, after Agibot helps Unitree, it leaves silently, and steadily puts the last piece of clothing into the washing machine (I guess it judges that the clothes draped over the washing machine are dirty), and this 10-minute one-shot footage finally comes to an end.
Although I don't know the architecture of this model, I believe everyone is as shocked as I am.
The first cross-body collaboration, the first time seeing autonomous evolution of robots, the first human-like autonomous decision-making to choose the optimal path, extreme spatial perception and force control, arbitrary task interruption, learning to use tools... Every movement and every frame of this model is an excellent demo, but they appear quietly in this 10-minute one-shot video, and the presentation is so simple and direct.
It seems to show a sense of dominance that these capabilities can be achieved easily and effortlessly.
The underlying technology breaks out of the inherent framework of mainstream models
Why can the robots in the video present so many iconic scenes?
I learned from insiders that the most important reason is that this set of models breaks out of all the inherent frameworks of current mainstream embodied models.
At this stage, mainstream embodied models can be roughly divided into three categories: VLA, WAM and traditional world models, all of which have bottlenecks that cannot be broken through for the time being.