HomeArticle

Models are evolving at a dizzying speed, yet the supporting evaluation systems cannot keep pace? The "competencies" of embodied intelligence are in urgent need of a unified "benchmark".

36氪的朋友们2026-08-19 09:18
Multiple parties have launched embodied intelligence evaluation platforms, and the industry is accelerating the exploration of unified standards.

The multi-stakeholder co-built regular embodied intelligence simulation evaluation platform RoboColiseum is officially launched, aiming to establish a simulation evaluation system for embodied models with credible results, reproducible processes, and high consistency with real machine performance.

More than half of 2026 has passed, and embodied intelligence models are emerging rapidly. However, the growth in the number of models has not made the industry clearer. On the contrary, the more optional solutions there are, the less unified reference there is for horizontal capability benchmarking.

Over the past year, multiple institutions at home and abroad have successively released embodied foundation models and VLA (Vision-Language-Action models), and the operations such as grasping, placing, pouring water, and folding clothes in demonstration videos have become increasingly "smooth".

However, the current embodied intelligence industry has always faced a practical dilemma: models are developing at a fast pace, but evaluation cannot keep up. The whole industry still lacks a systematic and credible measurement scale to clarify the real capability boundary of models and where their shortcomings are concentrated.

*Kechuangban Daily* has learned that the regular embodied intelligence simulation evaluation platform RoboColiseum, co-built by universities, scientific research institutions, open source communities, robotics enterprises and model companies, was officially launched a few days ago, aiming to build a simulation evaluation system for embodied models with credible results, reproducible processes, and high consistency with real machine performance.

Wu Mo, project lead of RoboColiseum, disclosed two core verification data in an interview with *Kechuangban Daily*: the correlation between the platform's simulation evaluation and real machine deployment reaches 89.5%, and the evaluation difference of the same model in the simulation environment and the real world is less than 10%.

This conclusion does not come out of nowhere. According to Wu Mo, the team carried out 50 to 100 controlled experiments for each of more than a dozen evaluation tasks ranging from simple to complex in both real machine and simulation environments, and finally obtained the 89.5% correlation through linear fitting.

RoboColiseum Platform Sim2Real Experiment Comparison

Different from the traditional evaluation method that summarizes model capabilities with a single comprehensive success rate, RoboColiseum sets multiple task sets around different capability dimensions to examine the model's performance in each dimension. It is understood that RoboColiseum has launched 4 sub-lists of capabilities including instruction following, spatial understanding, disturbance adaptation and general operation, as well as 78 high-fidelity simulation evaluation tasks, covering more than 3000 generalized cases.

"We summed up the four capability sub-lists based on examining what different dimensions of capabilities a robot needs to complete a series of tasks in the real world," Wu Mo explained the design logic to the reporter, "This evaluation dimension will test whether the model can understand instructions, understand the spatial world, resist disturbances in the real world, and finally complete the relevant tasks of instruction operations."

At the same time, the platform also disassembles each task into sub-steps to realize traceability of failure causes. During the evaluation process, the system not only determines the final success or failure of the task, but also records the steps completed by the model, failure nodes, and its generalization performance in different scenarios throughout the whole process.

Since the internal test, more than 40 teams at home and abroad have carried out model training and evaluation on RoboColiseum. Wu Mo revealed that generalization, long-range fine operation, etc. are the shortcomings of most current models, and relevant evaluation dimensions will be supplemented and updated to the platform later.

In fact, RoboColiseum is not the only one trying to make up for the shortcomings of evaluation standards. Since the beginning of this year, multiple forces ranging from official institutions to enterprises and academic communities are accelerating the standardization of embodied intelligence evaluation standards.

On June 1, 2026, YD/T 6770-2026 Artificial Intelligence - Key Basic Technologies - Embodied Intelligence Benchmark Test Method, approved and issued by the Ministry of Industry and Information Technology, was officially implemented, which is the first industry standard in the field of embodied intelligence. According to the introduction, this standard regulates the environment setup, task library construction, test process and index calculation method for carrying out embodied intelligence benchmark tests in simulation environments and real environments.

While the industry standard is being implemented, many enterprises have also launched evaluation platforms. Guanglun Intelligence officially released RoboFinals at the Beijing Artificial Intelligence Innovation Highland Construction Promotion Conference, building a standardized evaluation system based on 100 high-difficulty tasks; Dexmal and Hugging Face jointly launched the open benchmark test platform RoboChallenge, and the total number of real machine tests executed on the platform has exceeded 40,000.

Internationally, the University of California, Berkeley, Stanford University, NVIDIA and other institutions jointly launched an international public evaluation platform for embodied intelligent robot policies. The platform adopts a distributed, crowdsourced real-world evaluation network, and verifies the effectiveness of the method through more than 600 double-blind real robot evaluation comparisons.

Although all parties are accelerating their entry into this field, there is still a long way to go before the embodied intelligence evaluation field forms a unified benchmark like MMLU and GSM8K in the large model field.

Why is embodied evaluation so difficult? The essential difference lies in the different attributes of the evaluation objects. The evaluation of large models is mainly based on static question-and-answer datasets, and scores can be compared after one run. However, the evaluation of embodied models needs to be carried out in the dynamic physical world, which is difficult to compress into a set of standard questions and answers. Although real machine testing is the most reliable method, its cost is extremely high: a humanoid robot costs hundreds of thousands of yuan at every turn, and supporting test sites, operation and maintenance engineers, and long-term debugging cycles are also required.

In addition, the biggest pain point of simulation evaluation lies in the Sim2Real gap. Realistic factors such as light changes, object position offsets, and material differences may cause a model that performs well in simulation tests to fail to adapt to the actual deployment scenario.

Wu Mo also pointed out that there is no unified and authoritative evaluation list for models in the embodied industry. This is mainly because the industry and models are developing so fast that most evaluation benchmarks are no longer updated after being released, and cannot keep up with the iteration pace of models; at the same time, if the gap between simulation and real machines is too large, the evaluation results cannot truly reflect the model's performance.

Embodied intelligence is at a key node from laboratory demonstration to large-scale implementation. As the core infrastructure for industrial development, the evaluation system is becoming increasingly important. From the implementation of official standards to the successive entry of multiple platforms, the industry is accelerating exploration on the way to forming consensus, and *Kechuangban Daily* will continue to pay attention.

This article is from the WeChat official account "Kechuangban Daily", author: Li Jiayi, published with authorization from 36Kr.