HomeArticle

10.7 times faster, this edge-side inference engine has been open-sourced, and robot bodies can finally run large models without lagging.

机器之心2026-09-29 12:28
From RLinf to APXInf, Infimind has filled the last missing link for the implementation of embodied intelligence.

The second week of September was bustling in the embodied intelligence community.

On the 8th, Songyan Dynamics released the HERON-World Model; on the 9th, Agibot launched AGILE 2.0 and GE-Act 2.0 in one go; on the 10th, Unitree open-sourced UnifoLM-WLA-1.0 with 6B parameters, around 2500 hours of real robot data, and one model orchestrating 64 tasks; by the 15th, Songyan supplemented the release of HERON-CRA. In just eight days, three robot ontology vendors launched five models.

At the same time, robots are accelerating their march into the physical world. On September 20th, Qiyuan Robotics held a new product launch event, where two personal robots were officially released with the Q1 model priced starting from 19,999 yuan. Earlier on September 10th, UBTECH announced that it had secured overseas orders worth over 50 million yuan, covering products such as Walker C1 and UBTECH U1. The evaluation criteria of the capital market have also changed accordingly, and stakeholders are starting to re-examine embodied enterprises using metrics including de-watered repurchase rate, operating net cash flow, and actual fulfillment cost (i.e. the investment in deployment, debugging and maintenance after delivery).

When the two clues converge, a core pain point emerges: How to deploy the complex and powerful large embodied models efficiently and stably to the robot ontology hardware with extremely limited computing power? This is exactly the core value of the end-side inference engine.

On September 15th, Tsinghua University, in collaboration with InfiniteCor and Shanghai Jiao Tong University, open-sourced APXInf, an end-side inference engine tailored for embodied models. It answers two questions at the same time:

In the end-side environment with limited computing power, memory and power consumption, how to make the embodied model reach a usable inference speed on the robot ontology?

When the model keeps iterating and evolving, how to build end-side optimization capabilities that can continuously adapt and never fall behind?

For the first question, APXInf gives the answer with a set of figures: Without modifying the π0.5 model itself, through end-to-end full-stack optimization, the inference latency is reduced from 278ms to 26ms under the Thor chip FP8 configuration, with an end-to-end speedup of about 10.7 times, and the frequency of 38.46Hz brings robot control into the real-time range.

The answer to the second question lies in the construction method of the APXInf repository. It precipitates the model adaptation, optimization and verification that originally relied on a few experts into a set of continuously reusable workflows that can be used by Agents.

Project address: https://github.com/RLinf/APXinf-robo

Why does embodied intelligence require dedicated end-side inference optimization?

To answer this question, we must first clarify the position of the inference engine in a robot.

An embodied ontology is roughly composed of a main control unit, a computing power box and peripherals. A control loop goes like this: the main control unit collects observations, the computing power box infers the action chunk, sends it to the robotic arm or chassis for execution, and then moves to the next frame. The inference engine is embedded in the middle of this loop, which determines how many milliseconds an inference takes, what hertz the control frequency can reach, and how large, how hot, and how expensive the computing power box will be.

This position puts forward a set of very specific requirements for the engine: small batch, real-time, low latency jitter, and stable invocation by the main control unit via websocket or ROS. These are exactly the areas that cloud inference frameworks are not good at.

The approach of general inference frameworks is to lower the model layer by layer with a unified intermediate representation, and then hand it over to the backend to generate executable code, with one compilation stack covering as many models and hardware as possible; solutions like vLLM and SGLang are designed around cloud throughput. They are all successful in their own right, but their benefits are built on the premise of a wide variety of models, large batches, and sufficient scheduling space, none of which are met by the embodied end-side.

As a result, the end-side deployment of embodied models has formed three types of practical bottlenecks.

End-side performance. On the end side, computing power, bandwidth, power consumption and heat dissipation are all limited at the same time, while one inference needs to complete multi-view perception, model forward propagation and action generation, and respond to the main control unit at a stable rhythm. The memory bandwidth of end-side modules is 4 to 8 times lower than that of discrete graphics cards, and the power consumption is 5 to 10 times lower. Therefore, models that run smoothly on the cloud may perform poorly when ported to Thor or Orin.

Human resources and time. Deploying a model to end-side hardware is not a simple copy and run process, but a huge systematic project. It needs to go through the full-link process from underlying architecture adaptation, core operator compilation, precision quantization, to software and hardware performance tuning and simulation verification. This process usually takes several weeks and relies on cross-domain experts who understand both inference systems and operator optimization, and such talents are very scarce in any embodied company. More critically, minor changes in hardware can lead to all previous efforts being wasted: once the chip is replaced, all operator selection, memory layout and pipeline arrangement must be restarted from scratch.

Stability. A working demo and long-term stable operation are two different things. The latter has to face continuous perception control, limited resources and multi-module collaboration. The most feared problem is not slightly lower speed, but jitter, crash or state loss that occurs during operation.

The combination of these three points presents a multiplicative feature:

The total workload of deploying the model to the ontology ≈ (one-time access + one-time tuning) × number of ontology models × number of chip platforms × number of model iterations.

Every variable on the right is increasing sharply: the number of ontology models is growing, chips are becoming more diversified, and the model iteration cycle is also shortening rapidly. In fact, the five models launched in eight days mentioned at the beginning of the article are a direct manifestation of this situation. Expanding the team is not a realistic solution. What is needed is a layer of continuously reusable infrastructure: it can not only push single inference to the limit of hardware, but also make the access of the next model and the next chip no need to start from scratch.

What APXInf aims to build is exactly this layer.

What did APXInf do right to push inference into the real-time range?

Let's look at the first challenge first: how to maximize inference efficiency on limited hardware.

APXInf does not take a unified general IR as its primary goal, but prioritizes building specialized execution paths for specific model families. The model structure, weight layout, memory space, operator fusion scheme and execution order are all directly reflected in the code; only the common capabilities verified by multiple models will be further precipitated into shared modules. This design allows optimization to go deep into the model structure and hardware features, make targeted adjustments in operator selection, memory layout and execution process, and reduce the extra overhead introduced for compatibility with general scenarios.

At runtime, only the control capabilities truly required for end-side real-time inference are retained:

Operator execution uses CUDA Graph to complete full-graph capture and steady-state playback, and Kernel selection is generated and persisted by Autotune;

For memory, the model layer holds a fixed Workspace with fixed shape, pre-allocated and stable address, which reduces data movement on the hot path;

Scheduling is centered on small-batch real-time inference, and does not introduce Continuous Batching and Paged Attention designed for large-batch throughput.

No complex heuristics are performed at runtime, which makes the execution path predictable, reproducible and auditable.

This specialization is carried through to the construction phase: during compilation, the computing power of the local GPU will be queried, and the kernel will be compiled only for this specific architecture; the source code of CUDA kernel, CUTLASS and FlashAttention is built into the repository, no Docker or external framework dependencies are required during deployment, eliminating the need for images that are often several gigabytes in size.

Ultimately, this full-stack optimization can bring a huge improvement in efficiency. The optimization ladder on Orin can illustrate this point: the baseline is 1300ms, and after step-by-step optimization via torch.compile, Pipeline, Graph, Kernel, and Pruning, the latency finally drops to 119ms, with an overall improvement of more than 10.9 times. The same effect is achieved on Thor.

In terms of effect, the official evaluation protocol covers 50 episodes for all 10 tasks in LIBERO-10, with the seed fixed at 7, the replan step size set to 5, and a total of 500 rollouts. The success rate of Thor FP8 is 92.2%, Thor BF16 is 92.8%, Orin BF16 is 92.0%, and the reference implementation of π0.5 as the control group is 92.4%.

In addition to running fast, it also needs to run stably. The bottom layer of APXInf focuses on high-performance operators to squeeze the ultimate performance; the main part of the inference framework is developed using the Rust language. As a system language, Rust can provide low-cost, fine-grained system runtime control to achieve better scheduling and concurrency management. At the same time, Rust's mandatory RAII feature and ownership system greatly reduce the risk surface of memory problems, support safer and more robust resource lifecycle management, and the unsafe boundary is strictly limited to the fixed FFI boundary. The system-level value of this middle layer is to eliminate hidden faults such as wild pointers and data races as much as possible before launch; for a robot that needs to work continuously, it not only needs to "run at 38Hz", but also "still run at 38Hz after eight hours of continuous operation".

How can inference optimization keep up with the continuous iteration of models?

While the specialized path brings performance, it also brings new problems: manually writing a dedicated path for each model family means that the path has to be redone once the model is updated or the hardware is upgraded. If this work still relies on a few experts to complete manually, it will never be able to keep up with the iteration speed of embodied models.

APXInf's solution is to productize this set of engineering capabilities itself: it precipitates the model access, pre-processing and post-processing, ontology adaptation, performance tuning and deployment verification that were originally scattered in the experience of a few experts into engineering workflows that can be understood, called and continuously iterated by code Agents.

In this brand new workflow, the human-Agent division of labor presents a disruptive mode:

The Agent is responsible for running the process: read the PyTorch Reference, generate the execution Ledger, implement the static path of the model, perform operator-by-operator alignment and regression, and finally complete Autotune, documentation and iteration.

Humans are responsible for setting standards: define architecture boundaries and module responsibilities, Kernel contracts and safety specifications, acceptance lines for accuracy, performance and tasks, and decide when to extract common features into shared abstractions.

What is handed over to the Agent is no longer simple auxiliary work, but the implementation part that consumes the most expert time, and humans evolve into the "judges" and "decision-makers" who control the overall situation.

Also because the implementation can be handed over to the Agent, the rigor of the verification system has become the new core moat. APXInf has established three layers of guarantees:

Full-link layered verification: from single operator, Layer to the complete model, combined with cross-alignment of Eager and Graph paths, to ensure that every step is accurate and controllable.

Fail-closed principle: unsupported parameters or hardware will directly report errors, and never output wrong results silently.

Model family isolation: shared capabilities are only sunk at the Kernel layer, and any change must go through strict layered regression to lock down risks.

What does this mean for users?

For embodied enterprises, it transforms inference optimization from a one-time project into a sustainable engineering capability. When the self-developed model is updated to a new version, there is no need to queue up for expert resources again; when switching to a new chip, there is no need to discard all the tuning experience from the previous round; the team size is no longer a hard constraint on the access speed. More importantly, expert experience has evolved from personal know-how into reusable assets for the whole team.