Seamlessly fluid in simulations but prone to frequent "fumbling" in the real world: DreamSpace makes up the "physics lessons" for embodied intelligence, allowing robots to progress from imitation to true understanding.
With exclusive seed round support from Sunwise Technology, Memo Space rolls out dual-core physical AI products to advance the implementation of industrial service robots
When a human infant learns to grasp a cup, it goes through countless cycles of "fumbling, dropping, and trying again" that generate real physical feedback. Today's robots work in the opposite way: they have learned from tens of thousands of grasping videos, move smoothly in simulation environments, but frequently "slip" when they step into the real world. The whole industry is asking: what on earth is holding back the large-scale implementation of embodied intelligence? The truth is that almost all robots are "taking an exam without preparation" when facing the physical world.
At the 5th Global Digital Trade Expo, Hangzhou Memo Space Technology Co., Ltd. (Memo) has made up this "physics lesson" for embodied intelligence. It provides two core outputs: the world's first Physical-WAM (Physical-World Action Model) with physical perception capability, and the industry's first RoboTwin-phys physical drift evaluation benchmark. One is designed to enable robots to "understand physics" and gain human-like prediction capabilities, while the other verifies whether the robots have truly mastered such capabilities.
Memo Space focuses on the R&D and application of world models. It is incubated by the AVS Advanced Visual System Research Center led by Academician Wen Gao from the Advanced Institute of Information Technology of Peking University. Its core team has accumulated in-depth research experience in the fields of embodied intelligence, video codec, spatial computing and artificial intelligence, with both academic expertise and engineering implementation capabilities. In August 2026, Sunwise Technology (688507.SH), a leading physical AI enterprise, completed the exclusive strategic seed round investment in Memo Space. The two parties carry out in-depth collaboration in data, computing power, physical AI and other technical dimensions, to jointly promote the R&D of world models and their industrial implementation. Physical-WAM and RoboTwin-Phys are exactly the new answer sheet the two sides jointly present to the industry.
Hundreds of thousands of times of data gap creates huge barriers for the popularization of embodied intelligence
The year 2026 is a red-hot period for embodied intelligence. In the first half of the year, the total financing of China's embodied intelligence track reached 935 billion yuan, about 5 times higher than the same period last year. The shipment of humanoid robots exceeded 40,000 units, accounting for about 97% of the global market. China has become a critical hub for the manufacturing and shipment of humanoid robots across the world.
However, behind the booming trend lies the long-standing paradox of "hot capital but difficult implementation". One investor put it bluntly: "The order growth curve can never catch up with the financing growth curve."
Li Minghao, CEO of Memo Space, delivered a keynote speech titled *Physical-WAM: A Physical-World Action Model Trained on Observable Data and Deep Physical Data* at the Data Element Governance and Marketization Exchange Event of the Global Digital Trade Expo, which clearly explained the root cause of this challenge: the lack of basic cognition of the physical world for robots is the result of the superposition of two dilemmas.
The first dilemma comes from the data gap. The stock of real physical interaction samples for robots is only at the million level, which is hundreds of thousands of times less than the tens of billions of training materials such as text, images and videos. Collecting physical interaction samples through real robots costs a huge amount of resources, so the vast majority of models can only learn the mapping relationship between visual images and actions, just like students who only memorize by rote. Essentially, they are only imitating the actions they have seen, without understanding the physical laws behind them.
The second dilemma is the absence of evaluation criteria. Most of the existing evaluation systems only focus on the task completion degree at the visual level, ignoring invisible physical variables such as friction, mass and stiffness. For example, under the same thrust, a can with high friction will topple over in place, while a can with low friction will slide away directly. If you hold the rolling pin at the same position, the offset of the center of gravity will cause the rod to swing. These common "physical drift" phenomena in the real world can hardly be effectively verified by traditional evaluation systems. If a capability cannot be measured, the iterative optimization of the model will lose its anchor.
Two dilemmas correspond to two solutions. The solution proposed by Memo Space is a brand-new technical architecture.
Physical-WAM: Enabling embodied intelligence to evolve from "imitation" to "understanding"
Physical-WAM, the first solution launched by Memo Space, is the industry's first WAM base model with physical perception capability, which is trained on observable data and deep physical data. Its core difference from traditional models lies in the fact that traditional models can deduce the evolution of the physical world but cannot perceive physics itself -- they know "what the next frame will look like" but do not know "why it becomes like that". In contrast, Physical-WAM can perceive physical changes in real time during real tasks, predict the consequences of actions, and modify behaviors in real time.
This capability is realized through the collaboration of three modules. PhysLens acts as a "physical translator", extracting unobservable physical information such as friction, weight, center of gravity and contact state from multi-modal perception signals, and outputting a unified physical token representation. PhysDream works as a "physical prophet", combining the current physical representation, environmental state and candidate actions to deduce the future changes of physical states, predict risks such as slipping and falling off, and provide uncertainty assessment. Finally, PhysAct is responsible for physical-conditional action output, introducing physical attributes and future state prediction into action generation, and automatically adjusting the execution strategy when physical drift occurs in the environment.
The three modules form a complete closed loop of "perception → prediction → generation → modification". To give an intuitive example: when a robot grasps a paper cup, PhysLens identifies the low-stiffness feature of the paper cup material, PhysDream predicts that the increased load after water injection may lead to cup wall deformation or slipping, and PhysAct adjusts the grasping force and contact position accordingly, continuously modifying the strategy during the water injection process until the operation is completed. This is no longer a simple action reproduction, but a flexible response based on "physical intuition" just like humans do.
RoboTwin-Phys: Establishing a measurement standard for "physical intelligence"
Beyond models, the second solution provided by Memo Space is an evaluation criterion.
This evaluation benchmark follows a three-level progressive logic: whether the physical attribute inference is accurate, whether the task can still be completed under physical disturbance, and what the upper limit of performance gap is. Based on RoboTwin, it adds 13 types of physical disturbance factors common in the real world, covering 5 dimensions including object attributes, contact attributes, damping, task environment and camera configuration, and supports evaluation from single-factor disturbance to multi-factor combined disturbance. Its core value lies in converting abstract physical parameters in reality into controllable, markable and reproducible variables, so that "physical intelligence" has a quantifiable measurement standard.
How far is the "GPT Moment" for embodied intelligence?
Embodied intelligence and world models have not yet ushered in the "GPT Moment" similar to that of large language models.
Wang Xingxing, founder of Unitree Robotics, once put forward a measurement standard: when a robot is brought to any unfamiliar environment and can independently complete about 80% of the tasks through voice or text instructions, the industry will reach the critical point of explosive growth. "It may take 2 to 3 years at the fastest, or 5 to even 10 years at the slowest." Wang He, founder of Galaxy Universal, gave a more straightforward judgment: today's large embodied models are roughly at the GPT-2 level, and are expected to reach the threshold of GPT-3.5 to GPT-4 by 2028.
Behind the divergence of timetables lies the same gap. Industry experts point out that the existing several types of world models still have a long way to go to become base models that can truly understand, predict and interact with the real physical world. In other words, the distance to the "GPT Moment" is largely the distance between the model and the physical world.
At present, the world models in the industry are gradually divided into four categories: representational world models focus on abstract deduction of the world, generative world models focus on pixel-level reproduction of the world, interactive world models focus on 3D interactive simulation of the world, and the integrated world model of understanding, generation and prediction focuses on generating physical cognition to realize free manipulation of the embodied entity.
Memo Space is taking the direction of "integration of understanding, generation and prediction". Deducing the world is only a means, and the goal is to enable robots to understand the physical world like humans and adjust their action strategies in real time. At present, the company has formed a trinity of industrial collaborative ecosystem covering "data, model and evaluation". From the perspective of industrial implementation rhythm, it is expected to take the lead in exploring the joint verification of physical data, model capabilities and real scenarios in scenarios such as industrial operation, logistics handling and service robots.
Going forward, Memo Space will continue to iterate the multi-modal physical representation model along the path of "vision first, followed by force, touch and sound", expand the physical interaction dataset, optimize model inference latency and multi-entity adaptation capabilities, and collaborate with upstream and downstream partners in the industrial chain to complete the three key pieces of puzzle: physical data, world models and evaluation standards.
"Only when robots truly understand the physical world can embodied intelligence be widely deployed from laboratories to real industrial scenarios," said Li Minghao. Memo Space will adopt an open source and open attitude to promote the whole industry to accelerate the arrival of the "GPT Moment" for embodied intelligence.