A Tsinghua-affiliated embodied intelligence enterprise has secured financing of hundreds of millions of yuan, and its robots have been deployed on a large scale in JD Supermarket | HardKr Exclusive First Release
Author | Huang Nan
Editor | Yuan Silai
Hard Krypton learned recently that embodied intelligence infrastructure supplier "Lingyu Intelligence" has completed hundreds of millions of yuan in Pre-A round financing. This round of financing was jointly invested by Zhuzhou Industrial Investment, Flyto Venture Capital, Future Frontier Venture Capital, Xinneng Venture Capital and other institutions. The funds will be mainly invested in expanding the production capacity of mass-produced robots, continuously upgrading the real machine data collection pipeline, and iteratively developing the cloud operation platform. Maple Pledge Capital has long served as the private equity financing consultant for the company.
This is also the third round of financing completed by the company this year. Previous investors include Inno Fund, Meridian Capital, Galaxy Innovation Capital, Futian Capital, Leaguer Venture Capital, Tianying Capital, etc., with cumulative financing reaching hundreds of millions of yuan.
In 2026, embodied intelligence is moving from technical demonstration to the deep water zone of commercial implementation. The market's betting logic is changing. The growth of start-ups is no longer the only criterion, and what is being scrutinized more is the commercial feasibility of the path of "real scene data-driven model evolution".
Lingyu Intelligence was established in early 2025. The company focuses on the construction of underlying infrastructure for embodied intelligence, covering robot R&D, real machine data collection and cloud operation platform construction. At present, it has realized large-scale commercial deployment for open retail scenarios, and built a complete closed loop from hardware ontology to data pipeline.
The founding team has both academic background and industrial vision. Co-founder and Chief Scientist Mo Yilin is a tenured associate professor in the Department of Automation of Tsinghua University, studied under Richard M. Murray, academician of the National Academy of Engineering and pioneer of robot manipulation in the United States, with more than 10,000 citations on Google Scholar. Co-founder and CEO Jin Ge holds a bachelor's degree from the Department of Automation of Tsinghua University and an MBA from the School of Economics and Management of Tsinghua University. He used to be the managing partner of Yuanjing Venture Capital and vice president of Aoliang Photonics, and has many years of compound experience in venture capital and enterprise management in the high-tech field.
In addition, most of the core management team of the company graduated from Tsinghua University with bachelor's degrees, and the members have more than 20 years of working and trust foundation with each other. The main members of the team all have working experience in first-tier technology enterprises such as ByteDance, Kuaishou, Tencent and Meituan, and have accumulated rich industrial experience.
Lingyu TA series robots in home scenarios (Source / Enterprise)
From the perspective of the Lingyu Intelligence team, the underlying capabilities output by embodied intelligence Infra can be implemented as a deliverable system, which consists of mass-producible robot bodies, data assets with physical interaction labels, and a cloud platform that connects machines and models. The three support each other and form a complete closed loop.
This judgment is based on a premise repeatedly verified by the industry: if you want robots to obtain reliable operation capabilities in the physical world, the interactive data from the robot body in real scenarios is the key base to bridge the simulation-to-reality migration gap. Simulation data can undertake pre-training and sample expansion, but it is difficult to completely reproduce the infinite physical disturbances and long-tail edge cases in reality.
Mo Yilin, co-founder and chief scientist, told Hard Krypton, "The complexity of the physical world is beyond imagination, and there are many things you can't even think of." The premise of coping with this complexity is a truly "durable" robot. "It needs to be able to work continuously in real scenarios without failure, have sufficient force control accuracy to complete fine operations, ensure the consistency of long-term operation, and make every frame of data reusable."
Lingyu's self-developed TA series robots define the ontology based on this demand. Its core indicator is not "whether it looks like a human", but "stably completing tasks and outputting high-quality data". Based on this logic, the R&D team made trade-offs in several key dimensions.
Cost is the primary consideration. The price of Lingyu TA series robots is controlled at less than 1/2 or even 1/3 of similar products. Self-developed algorithms are used to replace expensive parts, and redundant configurations are eliminated. For example, the current feedback algorithm replaces the six-dimensional force sensor. The team uses a wheeled chassis plus a gripper solution instead of bipedal robots and dexterous hands. The latter contributes limited to data collection, but greatly increases system complexity and energy consumption. The hardware solution is lighter, and deployment and operation are also simplified accordingly.
Accuracy is the second key dimension. It not only refers to the positioning accuracy of the robotic arm, but also refers to the time and space alignment capability of the data link. At present, the TA series can achieve sub-millisecond time synchronization among vision, joints, force perception and control instructions.
Stability determines the "survival capability" of this solution in real scenarios. The TA series is designed around "continuous operation" from joint module selection, heat dissipation management to closed-loop control, ensuring that it maintains a repeatable working state even in environments with crowd occlusion and network fluctuations such as supermarkets, avoiding problems such as joint heating, vibration accumulation, and accuracy drift caused by long-term operation that affect data quality.
Lingyu TA series robots in warehousing and logistics scenarios (Source / Enterprise)
In terms of operation and control, it is the last checkpoint for the whole system to be "ready to use", and the core is to solve the end-to-end delay problem. The Lingyu team disassembled the entire operation link into more than 20 links and optimized them one by one, compressing the image transmission delay + force/position hybrid control response to less than 90 milliseconds. This means that the action synchronization between the operator and the robot is basically imperceptible to the naked eye.
In cross-city scenarios, the picture is stable, the operation is responsive, and task execution is not affected by geographical distance. At present, this capability has been continuously verified in real cross-city remote operations. After the TA robot enters the customer site, ordinary employees can operate independently after short training by professional engineers.
With the mass-producible robot body, the data collection pipeline starts. Since June 2026, Lingyu's data production capacity has risen from several hours a day to hundreds of hours, and it plans to reach a daily production capacity of thousands of hours by the end of the year. At present, the world's largest open-source real machine data set is on the order of 100,000 hours. According to the actual measurement of the Lingyu team, its database expansion speed can produce data of the same scale in about one month.
But data volume does not equal data value. The real differentiation of Lingyu at the data level is that every real machine data has complete physical interaction labels, such as force perception information such as joint torque, contact force, and end force curve. This is exactly the missing dimension in most current data solutions.
Lingyu Intelligence DexUMI (Source / Enterprise)
There is a deviation between the physics engine in the simulation environment and the real world. Human videos only record vision and cannot restore force and contact feedback. This information is exactly what robots need to complete operations in the real world, such as pushing open a door with resistance, taking an object with oily surface from a shelf, and maintaining grasping stability in a shaking environment.
Force perception data can only be generated when a real robot performs physical operations. Simulation, third-party video materials, and controlled Demo scenarios in the laboratory are difficult to restore the rich and varied contact feedback in the real environment, and it is difficult to produce high-quality native force perception interaction samples.
In this regard, Mo Yilin has a judgment, "As the scale of data expands, the industry will gradually transition from the pursuit of 'just usable' to the pursuit of higher quality data. When the model capability approaches the ceiling, data sets that lack multi-dimensional information will fail first."
The advantage of Lingyu is that its robots have achieved hundreds of units of large-scale deployment and work around the clock in real retail scenarios such as JD 7Fresh Supermarket. Every grasping, delivery, and physical contact with the shelf is synchronously recording force perception data, which flows back to the cloud to form a growing physical interaction database.
Based on this large-scale operating data collection system, Jin Ge, co-founder and CEO of Lingyu Intelligence, revealed that Lingyu Intelligence plans to build a million-level high-quality real machine data set within one year. The data set covers full time synchronization, space calibration and force perception labels, meets the engineering requirements for direct model training, and provides underlying support for embodied intelligence model training, scenario generalization and subsequent system evolution.
In addition to the robot body and data, Lingyu's iterating Nexus platform acts as the connection layer — connecting machines, operators and models together.
Currently, in the operation of JD 7Fresh Supermarket, Lingyu adopts the remote control mode. Operators face the screen in the command center and control robots in different stores hundreds of kilometers away.
Compared with the closed warehousing environment, real retail scenarios such as 7Fresh Supermarket pose completely different challenges to robots: the flow of people is dense and the moving lines are random, and customers may pass from any direction at any time; the goods on the shelves are constantly taken away, put back, and restocked, and the status is constantly changing; the timing and method of tasting and delivery need to be flexibly adjusted according to the stay and movement of customers. This means that every action of the robot cannot be fully rehearsed, and it must have the ability to work stably in a highly uncertain environment.
Lingyu TA series robots working in Taiyuan 7Fresh Supermarket (Source / Enterprise)
But the long-term positioning of Nexus is more than just a "remote console". It is designed as a human-machine hybrid collaborative platform. Simple tasks are completed autonomously by the machine, and when complex tasks or model confidence drops, cloud operators take over seamlessly. Every manual takeover itself generates "demonstration data", that is, the operator's precise control, force control, and response strategies to abnormal situations will be recorded, filtered and labeled for model iteration.
"This model itself is already creating value. Retail customers do not need to equip technicians in each store, operators can be managed centrally, labor costs are controllable, and service quality is standardized." Jin Ge told Hard Krypton. From one person controlling one machine to one person controlling multiple machines, and then to multi-person and multi-model collaboration, this is the evolution path set by Nexus. Jin Ge analogizes it to the gradual process of autonomous driving, "At the beginning, three cars were equipped with one safety officer, and now 50 cars can be equipped with one safety officer. We believe that only when robots really run in real scenarios, can their autonomy be gradually improved."
Back to the most fundamental question: what does an embodied intelligence infrastructure company ultimately need to solve?
In Lingyu Intelligence's narrative, the answer is clear and restrained. It is not to build an almighty general-purpose robot, but to provide the industry with scalable robot bodies, high-quality real machine data, and communication architectures that connect machines and the cloud.
From reducing the cost of a single piece of data to a daily production capacity of thousands of hours by the end of the year, from the single-point verification of JD Taiyuan 7Fresh Supermarket to the large-scale deployment in multiple cities such as Beijing and Tianjin, this path may not be sufficiently eye-catching. When more and more embodied intelligence companies need high-quality data to feed algorithms, Lingyu's role as a "data collection mother machine" will become an indispensable key puzzle in the ecosystem.
The following is an excerpt of the interview between Hard Krypton and Mo Yilin, co-founder and chief scientist of Lingyu Intelligence, and Jin Ge, co-founder and CEO (slightly edited):
Hard Krypton: The industry generally pursues fully autonomous operation of equipment, but problems occur frequently after entering the landing scenarios. How do you view the real capability boundary of current embodied robots?
Jin Ge: Full autonomy is indeed the final goal pursued by the industry, but from the perspective of industrial landing, there is still a clear distance between the realization of normalized full autonomous operation in highly open commercial scenarios and the ideal state.
In our customer docking, we noticed a phenomenon: many enterprises initially showed a clear preference for full autonomous solutions, but as they got closer to the actual scenarios, most of the full autonomous capabilities on the market remained at the level of single-point function demonstration. Once entering the real supermarket environment, after the superposition of high-disturbance factors such as the randomness of passenger flow, the variability of interaction methods, and the dynamic adjustment of shelf display, the system is still difficult to support stable operation in terms of reliability and generalization capability.
Hard Krypton: What are the underlying technical influencing factors behind the capability gap between "demonstration level" and "commercial level"?
Mo Yilin: The core lies in the generalization shortboard of model OOD (out-of-distribution) and its dependence on the training distribution.
At present, the capability of embodied models is highly dependent on the distribution of training data. For scenarios, objects and interaction logic that have been seen in the training set, the model can complete perception and operation tasks with extremely high accuracy. But when facing unseen edge scenarios and unknown disturbances, the fault tolerance capability will drop sharply. This is a typical "taking shortcuts" problem of AI models, that is, the model often does not really understand the essential logic of the task, but only fits the shallow features in the training data, which leads to the direct failure of the model as soon as the scene variables shift slightly.
A typical case is that when performing sorting tasks in a supermarket, the robot is asked to put red apples on a plate, and it performs stably under a clean and regular preset background. But once someone wears a red dress on site, the model will directly fail to identify and report an error, because the model does not establish the semantic representation of "apple", but only learns the feature mapping of "red = target".
The fundamental crux of this kind of problem is that it cannot be fundamentally solved by data augmentation in the laboratory environment. The artificially constructed scene variables are always only a partial mapping of the real physical world. The combination of variables such as the color distribution of customers' clothing, the randomness of shelf display, the change of natural light, and the difference in physical properties of the operation plane is almost infinite, and most of these variables will not be considered at the model training stage.
This is why Lingyu chose to let the robots directly enter real scenarios to work first. Only in continuous commercial operations, those Corner Cases that cannot be preset at all will naturally be exposed, and every exposure is an opportunity for model evolution. Edge cases with high frequency of occurrence accumulate data quickly and are solved iteratively first. Low-frequency abnormal scenarios can be covered by humans, which will not affect the overall commercial delivery. This also explains why we adopt the scheduling method of "half an hour of human operation and half an hour of autonomous performance" — on the premise of ensuring customer value, let the machine continuously accumulate real scene data within a controllable boundary to support the gradual evolution of the model.