36Kr Exclusive | Tsinghua-affiliated physical AI infrastructure startup raises tens of millions of US dollars in new round of financing, and its data equipment has entered the phase of large-scale global delivery
Source / Enterprise
This article totals approximately 3100 words, with a recommended reading time of 6 minutes
Author丨Ou Xue
Editor丨Yuan Silai
Hard Krypton has learned that Ropedia, a Physical AI data infrastructure service provider, has recently completed tens of millions of US dollars in angel round financing. This round of financing is led by a number of leading venture capital institutions deeply rooted in the Southeast Asian market, with multiple first-tier US dollar funds, smart manufacturing and large model industry participants jointly participating in the investment. In the previous round of financing, investors associated with Google, a16z, NVIDIA and Amazon also took part, with DeepVoyage Capital acting as the long-term exclusive financial advisor. The funds will be mainly used for the expansion of the core technical team, large-scale mass production of hardware product lines, and scaled production of data product lines.
Founded in the second half of 2025 and headquartered in Singapore with an office in Silicon Valley, the United States, Ropedia is one of the world's earliest companies to systematically build Physical AI data infrastructure. The company was co-founded by Chen Zhaoxi (CEO), Hong Fangzhou (CTO) and Liu Ziwei (Chief Scientist), associate professor at Nanyang Technological University, and its core team members are from Tsinghua University, MMLab of Nanyang Technological University, Meta Reality Labs and Google.
From self-developed collection hardware, data engines to large-scale real-world experience datasets, Ropedia is building the critical data layer that connects human experience and machine intelligence, providing a real-world data foundation for next-generation robots and embodied intelligence systems.
At present, the Physical AI industry is encountering a real-world data bottleneck. With the rapid iteration of robot foundation models, VLA and world models, traditional videos and image-text pairs can no longer meet the demand, and model training is in urgent need of high-quality data with real physical scale, dynamic interaction, hand-object relationship and scene semantics.
However, this type of data has long faced two major problems: collection relies on expensive equipment and complex deployment, leading to high costs; and there is still a long processing chain between original signals and trainable structured data. A piece of original collected data usually needs to go through processes such as time synchronization, spatial alignment, 3D and 4D reconstruction, human body and hand motion recovery, action segmentation, semantic annotation and quality control before it can be converted into training data that models can use directly.
Different from traditional data collection companies, Ropedia's core competitiveness is not hardware, but a complete data production system built around the Data Engine, which realizes model-driven data production to achieve large-scale supply of high-quality real-world experience data.
"Data is not necessarily better when it is more expensive. The key lies in cost, quality, scale, task relevance, and whether it can truly improve model performance and enter an evaluable training closed loop," said Chen Zhaoxi.
At present, most data collection companies on the market are still at the stage of single-point hardware sales or manual annotation services, while Ropedia can automatically convert massive original collected signals into structured Experience Data that can be directly used for model training and performance verification through its self-developed data engine Experience Engine, covering multi-dimensional information such as human posture, hand-object interaction, spatial structure and motion trajectory, which greatly improves data production efficiency and reduces data production costs while ensuring data quality.
Structure of Ropedia's multimodal Experience Engine (Source / Enterprise)
In Chen Zhaoxi's view, data and models essentially describe the same real-world distribution: data records the world discretely through real samples, while models approximate this distribution through continuous functions. Therefore, the competition of Physical AI in the future is not only the competition of model capabilities, but also the competition of data engine capabilities.
Liu Ziwei, co-founder and chief scientist, compares this transformation to an energy structure shift: "Large language models are somewhat similar to fossil fuels: humans create language, record language, and deposit abundant language across the entire internet, but fossil fuels are limited after all. Multimodal is the only way for AI to move from virtual to real, which is similar to the process of us shifting from fossil fuels to new energy."
The company has launched the head-mounted portable collection system HOMIE, which can lightly collect multimodal signals such as human motion, scene changes and object interactions from the first-person perspective, and continuously provide real-world experience data input for the data engine, forming a data flywheel covering data collection, data production, model training and verification.
Participants wearing Ropedia's head-mounted portable collection system HOMIE (Source / Enterprise)
In Ropedia's product system, HOMIE is only the entry point. The company is building a three-layer end-to-end real-world experience infrastructure: the first layer Sense is perception hardware covering different interaction scenarios, including HOMIE and perception systems for hand movements, tactile and spatial motion; the second layer is the automated production engine Experience Engine that converts original sensor data into trainable data, all data is stored in the database after manual pre-screening and cleaning, automatic annotation, indicator quality inspection and final manual inspection, with the whole process traceable; the third layer Deliver is the dataset and model training verification system, which is used to judge the actual effect of data on VLA and world models, and reversely guide the next round of collection.
In addition, Hard Krypton has exclusively learned that Ropedia is advancing the development of its next-generation multimodal data collection system, which is scheduled to be officially released in August 2026.
According to Chen Zhaoxi, the core upgrade of the new generation system is not simply increasing the number of sensors, but further unifying multimodal information such as first-person vision, hand movements, human motion and spatial position to achieve a more complete representation of real-world data. Compared with traditional videos that can only record "what people did", the next-generation multimodal data will further describe the interaction process between humans and the environment, providing higher-quality training data for robot foundation models, world models and spatial intelligence systems.
"The end point of multimodal is not to have as many sensors as possible, but to make vision, movement, contact and spatial state exist in the same time and the same coordinate system, forming a physical process that models can learn," Chen Zhaoxi gave an example. Tasks such as plugging and unplugging, twisting, pressing, and assembly are highly dependent on contact information, and visually similar actions may correspond to completely different force states and task results.
Ropedia's next-generation data visualization demo (Source / Enterprise)
In terms of data security and compliance, Ropedia has also established a complete delivery system. According to the needs of different customers, the company provides multiple solutions including cloud deployment, exclusive partition deployment and even private deployment in the customer's local computer room, ensuring that structured delivery can be completed without the original collected data leaving the customer's environment. This is particularly important in the Physical AI field involving real human behavior data.
In March 2026, Ropedia released and open-sourced the human experience dataset Xperience-10M. This dataset includes approximately 10 million real-world interaction clips, more than 10 synchronized modalities, with a total size of nearly 1PB. This dataset has been incorporated into the research and training systems of embodied foundation models, motion prediction models and world models such as Qwen-VLA, MolmoMotion, ACE-Ego-0 and τ₀-WM. At present, it has ranked Top 2 in the total download volume of all datasets on Hugging Face, and Top 1 in the embodied intelligence dataset category (cumulative downloads have exceeded 2.7 million times), and has been adopted by the world's top AI laboratories such as AI2 and Qwen. The company now serves about 1800 data-using institutions, of which more than 500 have signed data use agreements, and its customers cover more than 20 leading global embodied intelligence enterprises and internet technology companies.
The value of data has begun to be verified by model training results: ACE-Ego-0 used 435.7 hours of first-person human videos from Xperience-10M, and its ablation experiment showed that the task success rate of the model using both robot data and human videos reached 72.8%, which is higher than the 68.3% of the model using only robot data. The 4.5 percentage point improvement comes directly from the training data itself.
Chen Zhaoxi revealed that the company's turnover has increased by about 4.1 times in the past six months, with the North American market contributing 60% to 70% of the revenue.
In Ropedia's view, internet data determines the upper limit of large language models, while real-world experience will determine the upper limit of Physical AI.
The following is an excerpt from the dialogue between Hard Krypton and Chen Zhaoxi, CEO of Ropedia (edited):
Hard Krypton: What kind of data does the market actually need at present? How does Ropedia define the value of data?
Chen Zhaoxi: Industry needs are hierarchical. The first layer is massive and diverse original video data, which is the foundation. But in the next step, what customers need is the physical world CoT similar to the Chain of Thought of language models, not reasoning on the linguistic level, but causal footnotes in physical space. For example, when making coffee, the model needs to understand how the three-dimensional space changes, how hands interact with objects, and what state transitions the objects go through. These fine-grained, full-modality structured data with causal relationships are what the industry really needs in the next step.
The value of data depends not only on quantity and cost, but also on whether it can make up for the shortcomings of the model and improve its generalization ability.
Hard Krypton: Data collection cost has always been an industry bottleneck. How did Ropedia reduce the cost by 50 times?
Chen Zhaoxi: Costs can be split into four parts: labor, hardware, computing power and storage, and transmission. Our cost reduction is also focused on these four parts. In terms of hardware, we use self-developed equipment to replace consumer-grade hardware, remove functions that are not related to data collection, and achieve extreme cost compression through supply chain integration. At the level of computing power and storage, we have designed a sophisticated data pipeline, managing data versions like code to avoid repeated production. In terms of transmission, we have a self-developed compression protocol — for PB-level Physical AI data, logistics costs cannot be ignored. Overall, our collection cost is an order of magnitude lower than the industry average.
Hard Krypton: You judge that Physical AI will move from 0.1 to 1 in the next three years. What is the basis for this judgment?
Chen Zhaoxi: The core sign is the transition from PoC to small-scale deployment — from several robots doing pilot tests in factories to 100 or 200 robots truly participating in the production process 7×24 hours, running continuously for more than three months and being irreplaceable.
This judgment is based on several observations: First, the trend of large models extending to the physical world after their capabilities mature is clear; second, underlying technologies such as robot hardware, 3D vision, training and reasoning infrastructure have accumulated to a critical point; third, the willingness to pay and demand scenarios in downstream industries are rapidly converging, no longer scattered experiments, but productivity substitution with clear ROI. Market heat itself is also a signal. Capital, talents and industrial division of labor are beginning to gather, all of which are accelerating the process from 0.1 to 1.