Amid the explosive growth of the embodied intelligence track, Reconova's robots are already handling luggage at airports.
While robots are racing ahead in various tracks, an embodied intelligence company has blazed a unique path of its own, evolving from visual intelligence to the field of embodied intelligence.
On April 29, at the 3rd China Embodied Intelligence and Humanoid Robot Industry Conference, Reconova delivered a keynote speech on breaking through the scenario-based implementation of embodied intelligence.
This enterprise, which has been deeply engaged in the AI field for 14 years, officially sent a signal to the public: machines that can see the world are now ready to start working with their hands. In a track where everyone is talking about generality and scale, it aims to be a player focused on practical, down-to-earth implementation.
From Visual AI to Embodied Intelligence
Reconova was founded in 2012. Up to now, it has fully experienced two completely different eras of AI.
In the AI 1.0 era, the core proposition of technology was perception: how to enable machines to "see" images, identify objects and understand scenes. It was the golden decade for large-scale implementation of deep learning, as well as the fanatical decade for visual AI companies to rapidly expand their market presence.
Over the past ten years, the visual AI track has undergone a cruel round of market reshuffling. At its peak, there were thousands of domestic companies labeled as "AI vision", capital flooded in frantically, and valuation bubbles accumulated rapidly. Then came the long process of deleveraging: financing conditions tightened, commercialization could not be achieved for a long time, and homogeneous competition crushed profits. Around 2019, a large number of players fell into difficulties one after another, and it was no longer uncommon for former unicorns to be sold at a discount or even cease operations.
In the past visual AI field, security and finance were the two most concentrated competition arenas. In contrast, Reconova focused on relatively inconspicuous scenarios: passenger passage centered on civil aviation airports, commercial real estate dominated by shopping malls, and assisted safe driving for freight commercial vehicles.
From an external perspective, this is a somewhat restrained decision. But precisely because of this, in an industry with an extremely high elimination rate, Reconova has become one of the few visual AI companies that have survived from the small model era all the way to the large model era and still remain at the top of the industry.
The benefit of focus is that the moat gets deeper and deeper. According to Frost & Sullivan's data, in terms of 2024 revenue, Reconova ranks first in China's civil aviation enterprise visual intelligence product market, with a market share of 8.9%; its products have covered one third of domestic civil airports, and the coverage ratio among large hub airports with an annual passenger throughput of more than 10 million has reached two thirds. Behind this are hundreds of millions of scenario-specific trainings, in-depth understanding of all business links of civil aviation, and long-term customer relationships established with airport operators.
In the AI 2.0 era, the proposition of technology has changed. Large models not only bring improvements in perceptual capabilities, but also extend from understanding to action. This technical inflection point is exactly the timing for Reconova to take a step forward.
Zhan Donghui, founder and chairman of Reconova, welcomes this change: "Over the past 12 years, we have been making 'eyes' — perceiving and understanding the physical world through vision. But now we start to move forward, toward the direction of the brain and hands. On the basis of understanding the world, we begin to make some decisions and perform some executions to help people get things done."
This also means that Reconova is no longer just a visual intelligence company. It is shifting its technical focus from perception and cognition to decision-making and execution, forming a complete closed loop from "eyes" to "brain" and then to "limbs". In terms of product positioning, it is transforming into a provider of embodied intelligent products for commercial scenarios that perform complex operations. This is Reconova's new label, and also the specific track it has chosen in the popular embodied intelligence field.
The Real Moat in a Noisy Track
The most mainstream narrative of embodied intelligence at present is generality. The more scenarios a robot can adapt to, the more attractive its story is, and the greater its valuation space will be. Under this logic, companies that focus on vertical scenarios seem to be naturally in a weak position in narrative.
Zhan Donghui believes that general capability is the stage for platform-based companies, which requires scale, ecology, and first-mover data network effects. But the barriers of vertical scenarios are never built by piling up parameters. They come from specific scenarios, in-depth understanding of customer business processes, and the know-how accumulated after countless times of solving problems together with customers, which cannot be achieved by simply stacking computing power.
At the technical level, Reconova has built a competitiveness matrix composed of a three-layer architecture.
The first layer is the perception base. This is the direct transformation of 14 years of accumulation of visual algorithms: object recognition, spatial understanding, pose estimation, and real-time perception in unstructured environments.
The second layer is the decision-making layer, with VLA (Vision-Language-Action) large model as the core self-developed direction. Reconova is building a VLA model for vertical scenarios, which unifies visual perception, natural language understanding and robot motion planning in an end-to-end framework, enabling the robot to become an agent that understands scenario semantics, makes judgments according to context and generates corresponding action sequences. Compared with the general VLA model, Reconova further introduces force sense and tactile sense, making the robot's behavior decision-making closer to the human multi-dimensional information unified decision-making mechanism. Reconova names this innovation VTFLA.
The third layer is the execution layer, that is, self-developed execution components to complement the capabilities of "hands" and "body". No matter how strong the perception and decision-making are, they ultimately depend on the completion quality of physical actions. Reconova's self-development investment on the execution side solves the problem of reliable operation of robots in unstructured environments, namely grasping strategy, force control, and the adaptability of end effectors to different object shapes. This is an extremely high engineering threshold, and also the most difficult gap to cross from demonstration to mass production deployment.
On the judgment of the commercialization path of embodied intelligence, Zhan Donghui believes that complex and unstructured special scenarios will achieve commercial operation before general scenarios.
General robots face dual constraints of technology and cost. It still takes time to meet both conditions at the current stage: having sufficient generalization ability and reducing the unit cost below the procurement threshold acceptable to enterprise customers. In contrast, special robots deeply adapted to a single scenario can be fully optimized for known constraints in technology and have more commercial feasibility in cost structure.
The Underestimated Hard Bone
Civil aviation is the first entry point for Reconova to cut into embodied intelligence, and also the foundation with the deepest accumulation. The first implementation scenario Reconova found is baggage handling.
Baggage handling has always been one of the links with the most concentrated labor costs in the civil aviation industry. Difficulties in recruiting workers, high turnover rate, and large fluctuations in manual efficiency affected by weather and flight schedules are problems that have plagued airports for many years.
Making this scenario truly work well is far more difficult than it seems. Zhan Donghui said that the baggage handling area is a highly unstructured working environment, which almost gathers the most unfavorable conditions for robot deployment.
First of all, there is an extreme diversity of object shapes: there is no standard for checked baggage of passengers. Trolley cases, soft canvas bags, cartons, and oversized special-shaped items are often mixed in the same batch. Each piece of baggage faced by the robot is a new grasping challenge. It needs to figure out where to grasp, what kind of force to use to ensure that the baggage is stable and undamaged, and finally find the best position for stacking.
Secondly, the space itself is irregular: the transfer area under the terminal building is not designed for robots, the aisles are of different widths, and the gaps between equipment are tight, so the robot's motion path needs to be planned in real time.
Last but the most critical point, there is a high-density human-robot collaboration demand: in the civil aviation operation system, the accuracy and timeliness of baggage are directly related to the on-time rate of flights and passenger satisfaction. In order to complete the full transfer of all baggage of the flight with different specifications in a short time, human-robot collaborative operation is the best processing method at present. But parallel operation means that the two sides will have high-frequency spatial interleaving at close range, and any delay in perception or decision-making may cause safety risks.
This is exactly the reason why general robots cannot operate smoothly here at present. The strong generalization ability of general robots means that they are "usable" in many scenarios; but "usable" and stably available in harsh production environments are two completely different standards. At the same time, the current cost structure of general robots also determines that they cannot achieve an acceptable ROI in such labor replacement scenarios for the time being.
Reconova's answer is to develop an intelligent robot specially designed for the airport baggage handling scenario. At the 2025 International Airport Exhibition, in the simulated terminal transfer area, the Xiaoyi Baggage Handling Robot smoothly transferred pieces of baggage of different shapes from the end of the sorting system to the downstream baggage trailer, and completed palletizing efficiently, opening up one of the weakest links in the intelligence level of the civil aviation system.
One of the core designs is the industry's pioneering human-robot collaborative operation mode. On the premise of full communication with customers, it realizes seamless collaboration between robots and humans through engineering design, allowing people to work side by side with machines safely and naturally. Robots are responsible for high-frequency, heavy physical handling and palletizing, and humans intervene and supplement beyond the capability boundary of robots. Each side performs its own duties, and the overall efficiency is far higher than that of pure manual operation.
Zhan Donghui said that in the actual measurement of airport projects, the Xiaoyi Baggage Handling Robot has been able to significantly reduce the dependence on the number of labor and reduce the manual labor load. At the same time, it increases the system throughput by 30%, and the baggage damage rate drops to 0.12%, which will also become one of the driving forces for airport operators to make purchases.
At present, Reconova is carrying out actual measurements in multiple airports, and plans to officially achieve commercial implementation in the second half of this year. While expanding in the domestic market, Reconova also includes the civil aviation markets in Southeast Asia and the Middle East that have similar baggage handling pain points in its overseas expansion vision.
In this fast-growing embodied intelligence track, Reconova has chosen a more specific path: to do a difficult thing well, so that customers can see quantifiable value in actual business scenarios.
If you want to find a coordinate for Reconova in the current industrial landscape of embodied intelligence, it is neither a general robot company nor a visual AI company in the traditional sense, but an embodied intelligent product provider focused on handling complex scenarios and complex actions.
The boom of the robot track will eventually fade, but products that have been verified in harsh scenarios will not. Amid the noise, insisting on narrowing down and deepening the track is a choice that requires concentration. But it is precisely this choice that allows Reconova to occupy a truly scarce ecological niche in the noisiest boom of embodied intelligence, and become an embodied intelligence company worth looking forward to.