Wang Xingxing revealed the "Story of Unitree's Value Appreciation"
The day after the listing, Wang Xingxing publicly reviewed the mental journey of Unitree's "growth".
On August 20, Wang Xingxing appeared at the 2026 World Robot Conference and delivered a speech titled From Exhibit to Product: The Next Decade of the Humanoid Robot Industry, systematically reviewing Unitree's "leap" in the past nearly ten years from technology, products to implementation scenarios.
The origin of Unitree's story dates back to 2016. Wang Xingxing recalled that the company initially focused on quadruped robots, and only shifted its focus to humanoid robots in recent years; the G1 launched in 2024 has become a landmark product in the global humanoid robot field — the vast majority of humanoid robots on the market are either derived from it, or have extremely similar appearances.
But compared to "making it", he cares more about "putting it into use". Wang Xingxing revealed that in the past few years, Unitree has been promoting robots to "truly enter daily life and work in factories". In this process, Wang Xingxing emphasized the word "truly" many times. In other words, Unitree also believes that stage performance is not the core implementation scenario for humanoid robots.
Shifting from pre-recorded and choreographed actions to multi-modal end-to-end actions generated by real-time voice, this is Unitree's technical leap route in the past 10 years. Wang Xingxing admitted that the seemingly smooth demonstrations in the past were all pre-trained actions; while the capabilities demonstrated now are generated based on real-time voice.
"We hope that when real robot applications are implemented in the future, all their actions will be generated completely in real time, rather than programmed in advance."
The consensus of the entire embodied intelligence industry is that data supports the capability of robots for real-time action reasoning and execution.
In Wang Xingxing's judgment, "All current robots are AI-driven, and AI is basically data-driven. The amount of AI you get is proportional to the amount of data you have, especially high-quality data."
In terms of specific data types, pre-training mainly relies on massive real human data and Internet data, plus the measured data of robots in the physical world. The former has a larger volume, while the latter is rigidly needed. "We need to truly match the robot with the physical world."
Wang Xingxing revealed that the company's capital and manpower investment in AI models "should be the largest at present". The world model is a tortuous route: the company started to develop world models based on video generation in early 2020, but the project was once shelved at the end of the year due to unsatisfactory effects, until it re-invested heavily in this direction in 2025. This rhythm of "trying first, shelving, then making efforts again" reflects the repeated exploration from 0 to 1 on the technical route of an embodied intelligence company.
In addition to data, moving from the stage to the factory requires the generalization capability of the model behind the robot body. The objective situation is that today's robots can complete simple assembly tasks, but their efficiency is still lower than that of humans, and they need to be retrained for every new task.
"If you perform sufficient collection and training in a fixed scenario, the success rate can be close to 100%, but if the operated objects or the environment change slightly, the success rate will drop sharply," said Wang Xingxing.
Wang Xingxing believes that the insufficient alignment between the input and output of AI models and the real physical world leads to poor generalization capability. "Language models are purely digitally coded with almost no loss. But every input and output of a robot has deviation and loss."
As for the solution to this problem, Wang Xingxing's vision is "physical AI robot self-evolution" — using the most cutting-edge large model to drive an agent: it searches for the latest papers and open-source solutions online by itself, writes control codes, verifies them in the simulation environment first, and then deploys them to the physical robot; AI automatic scoring combined with human experience scoring forms a positive cycle.
Regarding the question that the public is very concerned about, "When will general-purpose robots truly enter households?", the critical point indicator Wang Xingxing set is that a robot, when brought to any unfamiliar environment, can complete about 80% of the tasks through voice commands, "that will be the critical breakout point of the real industry in the future".
He predicted that it will take as fast as two to three years, or as slow as five to ten years. At the moment of the AI boom, the evolution speed of the entire robot industry has been greatly increased, "it may be a little faster than I estimated", and this is a "brand new starting point".
The following is the full transcript of Wang Xingxing's speech (with deletions and adjustments without changing the original meaning)
Hello everyone. First of all, let's review the situation of our company: Unitree was founded in 2016, and by August this year, it has been almost exactly ten years. The company initially made quadruped robots, and has been working on humanoid robots in recent years.
One of our relatively successful products in recent years is the G1 launched in 2024. After its launch, this robot has basically become a landmark product around the world; most of the robot products we see now, or many robot products in use globally, are basically derived from it, or have very similar appearances.
A representative case at the beginning of this year is our Wu Bot — we collected dozens of classic Chinese kung fu actions in advance, performed AI training, and then presented them on the stage. The biggest highlight of this program is the combination of the world's top humanoid robot technology and traditional Chinese kung fu and lion dance culture. The video has spread very well overseas, with tens of billions of views, and everyone loves it very much. I think this is a very good example of the integration of technology and culture.
In addition, we also held a simple demonstration at the Temple of Heaven, with about 49 robots performing fully automatically on site, which was very shocking. The Temple of Heaven has a cultural history of hundreds of years, and the robots there create a feeling of traveling through time and space — the culture of the past hundreds of years is integrated with the latest technology at present.
In the past few years, we have been promoting robots to truly enter daily life and work in factories. In 2024, we cooperated with automobile factories to carry out implementation applications; we also deployed robots in our own factories for simple tests and implementation. Most of our company's AI teams are working on making robots truly work in households or factories.
Why haven't we promoted it on a large scale now? The main reason is that the current efficiency and generalization capability are not sufficient.
Although robots can do some simple assembly tasks, their efficiency is still lower than that of humans; and every time a new task comes, they need to be retrained, which makes the efficiency even lower. Therefore, we hope to improve the generalization capability of the technology before large-scale promotion — this is the biggest common bottleneck around the world at present.
In April this year, we set a new record: the maximum running speed of humanoid robots exceeded 10 meters per second. This robot is modified from H1, our first-generation humanoid robot in 2023. H1 is also our company's first humanoid robot, which is very classic; it can also be equipped with wheels later for more flexibility.
As you know, all current robots are AI-driven, and AI is basically data-driven — the amount of AI you get is proportional to the amount of data you have, especially high-quality data. We collect data by ourselves, and also cooperate with third parties to collect more data. The training data for real AI robots is mainly massive real human data or Internet data as pre-training data, plus part of the real machine data, which are two rigidly required parts. In terms of volume, real human data is more and more important; but the actual real machine data is also very critical, because we need to match the robot with the real physical world.
In May this year, we released the world's first mass-produced manned deformed mech GD01, which is more than 3 meters tall and weighs about 500 kilograms after carrying a person. It is a bit like an off-road vehicle, not for use in households or cities, but for outdoor hiking and transportation in complex environments; the most interesting thing is that it can deform into a four-legged mode, with better stability and extremely strong obstacle-surmounting capability.
A few months ago, we demonstrated automatic AI action generation based on multi-modal end-to-end. Many of the conventional action demonstrations you see are pre-trained; while these actions are generated based on real-time voice — because of this, there will be a few seconds of delay in the middle: after each sentence is finished, it needs to go through speech recognition, upload to the cloud for AI action generation, and then be sent to the robot after verification. But we hope that when the technology is truly implemented in the future, all actions will be generated completely in real time, rather than programmed in advance.
We also promote robots to truly work in conference rooms. The conference rooms in our company are often messy, so tidying up the conference room is a very rigidly needed task. What we mainly collect and train is: after the conference room is messed up arbitrarily, restore it to a clean and tidy state. It is fully realized end-to-end, with about seven or eight different tasks set, which can be automatically switched and executed directly by one model, and is anti-interference. We have been developing this feature recently, because scenarios with practical value and not too much difficulty such as conference room tidying are very meaningful.
On the basis of automatic voice generation, we have made a version with working capability — driven by voice, the actions can perform practical work. For example, after equipping with a multi-modal model, if you ask it to "bring up the blue medicine box on the third layer", it can recognize and execute the instruction, and is anti-interference, instead of performing fixed actions.
Some time ago, we released the latest generation of wheel-legged robot Go2-W, which is relatively lightweight with a self-weight of only about 20 kilograms, but has very strong load capacity and battery life, is basically rainproof, and can be used indoors and outdoors. We transplanted some humanoid robot technologies to quadruped robots, further improving the flexibility of quadruped robots, and their load capacity is also very strong. We very much hope to implement them in scenarios such as outdoor hiking.
We also upgraded last year's small and medium-sized humanoid robot R1, as well as consumer-end products such as A2, which have very high cost performance and are available for pre-sale or retail on JD.com and Taobao. Some time ago, we previewed a new robot, which took only more than three months to develop; it is only a preview now, not an official release, but its maximum speed has reached 12.65 meters per second, exceeding the fastest running speed in human history, and its high jump height has also basically exceeded the maximum jump height in human history.
In 2025, I myself and our company's products were listed in the TIME magazine's rankings.
Looking back on the development of embodied intelligence AI models in recent years, the main schools are currently VLA models and world models, and the world model is what you hear the most this year. Our investment in AI models has always been very large, which should be the direction with the largest capital and manpower investment in our company at present. We started to develop world models based on video generation in early 2020, but the effect was not satisfactory at the end of 2020, so the project was shelved for a while; last year, we vigorously developed the direction of world models based on video generation again.
Many people are curious about when general-purpose robots will truly enter daily life and households? The biggest problem is still the current global bottleneck — the generalization capability of embodied intelligence is not sufficient. For many AI models, if sufficient collection and training are performed in a fixed scenario, the success rate can be close to 100%; but if the operated objects or the environment change slightly, the success rate will drop sharply.
Therefore, I hope that in the future, as fast as two to three years, or as slow as five or ten years, the day will come when you can see: bring a robot to any unfamiliar environment, to your home, and it can basically complete about 80% of the tasks, then the critical moment of embodied intelligence will be almost realized, and that will be the real critical breakout point of the industry in the future. The evaluation indicator is very simple: bring a robot to about 80% of unfamiliar scenarios, and it can complete 80% of the tasks through voice and language instructions, which is a very important critical point.
Why is it the hardest to reach this point? Because the current input and output of AI models are not sufficiently aligned with real robots. If you ask it to move, assemble things together, or move an object to a certain position, the general direction is fine and it can perform roughly, but the success rate or deviation of the last few centimeters or the last few millimeters cannot be well matched with the real world. When I reach for an object, I am about to grasp it, but the last bit of tactile feedback and the last bit of error cannot be well corrected, and the success rate will drop sharply. This is the biggest bottleneck of embodied intelligence around the world at present — there is deviation between the input and output of AI models and the real world.
Why do robots have this problem, while language models or multi-modal models do not? Because language models are completely digitally coded, their input and output are strictly limited in the vector space, there is basically no large deviation, and there is no loss — there is no loss at all when the input and output are reversed. But for robots, there is deviation and loss every time input and output are performed, so the generalization capability and success rate are slightly worse. But I think this problem can definitely be solved in the next few years, it just takes a little time.
Let me introduce a project that we are promoting recently and just proposed yesterday. In the past few years, everyone has used AI for a lot of programming and development, but the real application of AI in the robotics field is still relatively insufficient.
What we are promoting is: directly enable physical AI robot models to self-evolve. Driven by the most cutting-edge top large models at present, we set some rules, experience, constraints and tools by ourselves, let it search for the latest papers, the best research results and various open-source schemes on the Internet by itself, then it programs by itself to generate its own robot control codes; after the codes are generated, they can run in the simulation environment, or call the large model state to control a physical robot.
If the codes are deployed to the physical robot for testing, then evaluation and scoring are carried out — part of this is realized by the AI model itself, and part of it can also involve human participation in scoring, because human experience is very important, and humans can easily judge the quality of results. Once this logic runs systematically, the development efficiency of the entire robot will be much higher, and the iteration speed will also be very fast; after running for a period of time, the self-evolution capability of robots can be truly improved.
(PPT) The lower left corner is a coding agent, which writes robot control codes by itself. At present, for many simple deep reinforcement learning algorithms, the effect of automatic programming by the coding program is already good. Put the programming result into the simulation environment for verification, and then put the trained code into the physical robot for testing; after the test is passed, perform AI model scoring and manual scoring for evaluation, confirm the effect, and then feed back to the coding agent to form a positive cycle. This is the simplest example, and the actual use will be more complex.
I think this is a project that is very worth doing globally at present and in the next few years. The reasons are very simple: first, with the improvement of the capability of the basic model every month and every year, the capability of this self-evolution cycle will be further improved, and it can keep up with the progress of global AI technology; second, it can use diverse data more efficiently — the efficiency of manual data usage is very low, and the self-cycle can make full use of various training data, real world data and human data; third, more physical machine deployments can be realized. The more deployments there are, the more test data and evaluation indicators there will be. The data utilization rate and growth rate are guaranteed, which can achieve large-scale data, and some skills can be accumulated every day or every month, instead of developing a skill today and wasting it tomorrow.
I think this is a very important thing — to truly use AI to greatly improve the development efficiency and evolution efficiency of robots. This is also a new opportunity to enable the self-evolution of physical AI robots, and this direction will be promoted by many people in the future.
Our company first participated in the World Robot Conference around 2017 and 2018, and we were deeply impressed. I think in the next ten years, at the new starting point at present, especially in the era of AI boom and high global attention, the evolution speed of the entire robot industry has been greatly improved, which may be a little faster than I estimated, and this is a brand new starting point, a brand new beginning.
Due to time constraints, the presentation is relatively brief. Thank you all, and please forgive me for any shortcomings.
This article is from the WeChat official account "Tencent Technology", author: Su Yang, editor: Xu Qingyang, published with authorization from 36Kr.