Infinite Vision Tech Debuts at WAIC: From Navigation to Data Acquisition, Completing the Critical Chain for the Implementation of Embodied Intelligence
In July 2026, WAIC is held in Shanghai.
Walking into the embodied intelligence exhibition area, a subtle atmosphere mixed with anxiety and fanaticism fills the entire pavilion. At WAIC two years ago, only 18 enterprises ventured into this field; last year, the number soared to 80; and this year, the number of exhibitors has quietly exceeded 200.
Behind the growth rate faster than Moore's Law, there are not only various end players shining brightly; with the architectural dividend of model algorithms reaching a phased bottleneck, data Infra supply chain players represented by Context Technology, which were previously hidden deep under the water, have also stepped into the spotlight.
I. How Data Becomes the Achilles' Heel of Embodied Intelligence
To understand the importance of data to embodied intelligence, autonomous driving is the best mirror.
If we break down embodied intelligence into different links, in fact, intelligent driving is only a specific subtask in the map of embodied intelligence. But even giants like Tesla and a large number of domestic new car-making forces, with millions of user data continuously uploaded to the cloud and nearly a decade of dedicated research, still remain cautious when facing various strange Corner Cases.
The scenarios faced by embodied intelligence are more than an order of magnitude more complex than those of autonomous driving:
The first layer is geometric dimension elevation from 2D to 3D: automobiles only need to handle movement on the X-Y plane, while the hand-eye coordination of robots is carried out in a completely open 3D space.
The second layer is exponential explosion of degrees of freedom: automobiles only have steering wheels, accelerators and brakes, while the degrees of freedom (DoF) of humanoid robots often reach dozens, which means the range of possible movements increases exponentially.
The third layer is severe data asymmetry: compared with autonomous driving, the number of robots that can provide high-quality data feedback is very small.
More critically, embodied intelligence does not require ordinary video data, but embodied experience with spatial structure and task logic.
"There are massive amounts of videos on the Internet, but for embodied intelligence, videos without reliable pose are just garbage materials," a relevant expert from Context Technology said bluntly, "2D maps only need to describe planes, while robots need to know their absolute position, pose, motion trajectory and task state in 3D space. Only data with stable trajectory, time synchronization, spatial alignment and quality assessment can be regarded as trainable silicon-based experience. And these require data such as SLAM/VIO, robust 6DoF trajectory, real-time online output and trajectory scoring."
Over the past year or more, Context Technology has precisely targeted this data desert, and won cooperation with a number of leading large manufacturers by relying on underlying components such as Insight spatial intelligent camera and TinyNav navigation algorithm library.
But this year, Context Technology found that the pain points of the industry are no longer just a single hardware or navigation algorithm. A set of data closed-loop infrastructure that can be collected at scale, verified and reused is the foundation for the rapid evolution of various embodied intelligent terminals.
Therefore, at this year's WAIC booth, this company labeled as physical AI infrastructure launched two representative products: Looper Embodied Intelligence Data Collection System and NavCore Embodied Intelligence Autonomous Navigation Brain.
They are precisely targeting the two major problems encountered by the embodied intelligence industry at the data level: how to mass produce high-quality data in the real world at scale, and how to achieve highly stable environment perception and autonomous navigation on the robot body side.
II. Looper Data Collection System: Equipping Robots with First-Person Learning Capability
On the battlefield of embodied intelligence data collection, there are currently three mainstream technical routes: teleoperation + VR, wearable multi-view, and fixed multi-camera studio.
The biggest advantage of teleoperation is that it can directly obtain robot motion data, which is suitable for training robot execution strategies. However, it usually relies on special equipment, with high collection cost and limited efficiency. At the same time, the data distribution is easily limited by the robot's body shape and operation mode.
Fixed multi-camera collection has strong controllability and can obtain high-quality, multi-view data, but it relies on specific sites and equipment, making it difficult to cover a large number of real tasks and long-tail scenarios in families, factories and commercial environments.
Considering the above background, Looper chose the Ego-centric multi-view solution, whose core concept is: let humans complete tasks from the robot's perspective, and convert the physical experience accumulated by humans for tens of thousands of years into digital assets that robots can understand without perception and at low cost.
In the industry's view, this perception method that is closest to the way humans complete tasks is also the route with the greatest potential to realize the scaling of real scene data at low cost.
In terms of configuration, Looper adopts the golden combination multi-view solution of head + both hands: the head Insight camera is responsible for recording first-person visual information, spatial environment and its own motion trajectory; the hand cameras are responsible for capturing close-up operation details, including hand movement, object interaction and local spatial changes; the backpack end provides edge computing, synchronous recording and data storage capabilities.
But this is by no means simply strapping several cameras to the body for video recording. "It is easy to collect data, but collecting valid data is the barrier. What Looper does is to collect valid data and high-quality data," Yan Qinrui, co-founder and Chief Operating Officer (COO) of Context Technology, gave us an example: in the process of data collection, spatial and temporal alignment between multiple cameras is a very critical industry difficulty. If the images, depth information, inertial data (IMU) and pose trajectories from different perspectives cannot maintain high-frequency synchronization in the same space-time coordinate system, the trained robots will have movement disorders.
To solve this problem, Context Technology chose the technical solution of directly completing the 6-degree-of-freedom trajectory calculation of VSLAM on the end side and outputting it.
In the past, for embodied intelligence data collection, the common practice in the industry was to collect data for a whole day, take it back and run 3D reconstruction offline on the server for several days, only to find that the data alignment failed and the whole batch was scrapped. With Looper's direct output of VSLAM on the end side, the spatial position of each frame of image can be known in real time on site, and unqualified data can be re-tested immediately.
Another major advantage of this method is that the images, IMU and 6DoF (six degrees of freedom) trajectories are naturally aligned on the end side, and the cost of subsequent 3D reconstruction and task segmentation drops sharply.
During the final retrieval, what customers get is not a bunch of disordered video files, but high-quality experience packages that can be accurately retrieved by spatial area, motion trajectory and task stage, which truly enables the original video to be converted into truly learnable and verifiable data.
It is reported that Context Technology is also planning to launch a new generation of four-camera large-field-of-view data collection headband, with the core optimization direction of lightweight wearing and unconstrained natural collection. The new solution will restore the human layered perception logic: use peripheral vision to take into account the global environment, lock the line of sight on the interacting object in front of you, and complete the upgrade of the robot data collection solution from the underlying perception.
III. NavCore Navigation Brain: How to Make Robots Move from Being Able to Walk to Walking Skillfully
If Looper solves the problem of data production, then what NavCore needs to solve is the autonomous navigation problem of robots.
In the past, embodied intelligence companies were like senior system integrators when it came to navigation:
Buy cameras from Manufacturer A, select computing power boards from Manufacturer B, purchase positioning algorithms from Team C, and then the R&D team stays up all night writing drivers, adjusting sensors and aligning timestamps.
As a result, trivial and tedious work such as frequent disconnection of wireless networks during on-site deployment, uneven power supply of multiple boards, and interface adaptation often consumes a lot of time for the R&D team.
In addition, in the entire link, a tiny delay of the camera will amplify the positioning error, the positioning error will directly pollute the global map, and the slow map update will lead to path planning that hits the wall. Any jitter in any link may directly lead to task failure.
Not to mention, compared with the past 2D navigation solutions, under the same resolution, the data volume of 3D navigation will increase by more than an order of magnitude, and it will also increase the computing pressure of point cloud processing, collision detection and path planning.
The difficulty of software algorithms has escalated, but the conditions faced by hardware are more stringent: in the past, 2D-based autonomous driving could be carried on a large car; while embodied intelligence needs to deploy 3D navigation capabilities to low-cost embedded platforms.
Facing the above problems, as the autonomous navigation brain of embodied intelligence, NavCore's approach is: large-scale integration and standardization. It can encapsulate complex robot navigation engineering into deployable products, so that customers do not need to build perception, computing power, mapping, positioning and navigation links from scratch.
In terms of hardware, it takes the NVIDIA Jetson series as the computing power base, and is deeply integrated into the world's mainstream developer ecosystem; at the same time, it integrates the self-developed Insight spatial intelligent camera, computing, communication and power management into a modular integrated design to reduce integration complexity.
It is worth mentioning that Context Technology's self-developed Insight spatial intelligent camera, traditional cameras only need to clearly capture the picture at a fixed position to solve 90% of the problems, but in the embodied intelligence scenario, it not only needs high image quality, but also a sufficiently large field of view, extremely low perception delay, reliable sensor synchronization, and the ability to continuously output stable position and pose under vibration, sharp turning, low light, strong light and dynamic occlusion.
This requirement determines that the lens, sensor, inertial measurement unit, chip and algorithm cannot be designed independently. To this end, Insight adopts a 188-degree ultra-wide field of view, industrial-grade inertial measurement unit and end-side computing power to complete visual positioning, depth calculation and spatial perception locally on the device, so that embodied intelligence can not only see, but also accurately understand complex spatial relationships.
In terms of the final software algorithm, Context Technology also uses algorithm pruning, computing optimization, hardware acceleration and model quantization compression to embed the spatial intelligence that originally required server-level computing power into the low-cost robot end side, which can operate stably for a long time.
In addition, considering that many robot projects in the past need to re-do a set of system integration for different bodies and different scenarios, including sensor selection, computing power adaptation, driver development, coordinate system calibration, algorithm deployment and on-site debugging, etc. This leads to a long robot development cycle and difficult replication of solutions, which is also an important bottleneck for the large-scale implementation of the industry.
NavCore not only has Context's powerful TinyNav visual navigation algorithm built-in, but also can provide standardized interface output based on ROS2. Whether it is humanoid, quadruped, AMR or UAV, or scenarios such as global spatial perception and multi-robot scheduling in smart parks, centimeter-level autonomous inspection and dynamic obstacle avoidance for industrial inspection, real-time navigation and task execution in complex port logistics environments, intelligent path guidance and immersive explanation for cultural tourism guidance, it can be quickly migrated and used out of the box.
Through the underlying reconstruction of the entire navigation link, NavCore outputs the most complex and error-prone integration work of the entire industry through standardized interfaces. The R&D team of embodied intelligence can naturally release their limited energy to the upper-level task intelligence and application scenarios to accelerate innovation.
IV. Epilogue
As the scenarios move from simple demos with fixed trajectories to real scenarios full of uncertainty, embodied intelligence is ushering in its coming-of-age ceremony.
In the next three to five years, whether the invisible links such as data closed loop, perception navigation and body control can be solidly built will become the new industry survival standard.
The two answers handed over by Context Technology at WAIC: Looper is to encode the labor experience that humans have accumulated for tens of thousands of years in the physical world into the digital world at low cost and on a large scale; NavCore draws a deterministic travel route for robots in the continuity and irreversibility of the physical world, which may be the key.
Only when the infinite human experience begins to feed the model continuously through standardized pipelines, and the engineering threshold of autonomous navigation is reduced to out-of-the-box availability, can embodied intelligence truly cross the key step from expensive demos to general labor.