Roundtable: Intelligence Moves to Reality, Industry Embraces New Development — How AI Enters the Real World | 36Kr 2026 Industry Future Conference
In 2026, industrial investment has entered a deep-water zone, where capital, technology and industry are accelerating their integration. The old investment logic no longer applies, and a new consensus is taking shape. The 2026 Industrial Future Conference focuses on opportunities in the new cycle, and jointly explores the future of the industry and the birth of the "Light of China". From September 9 to 10, the 2026 Industrial Future Conference hosted by 36Kr, with the theme of "Above Deep Waters, Resonate for New Growth", was held in Yizhuang, Beijing. Representatives from state-owned capital platforms, industrial investment funds, corporate CVCs, innovative enterprises, as well as experts and scholars gathered to focus on the industrialization of future industries such as quantum technology. The conference conducted in-depth discussions on cutting-edge technologies and industrial perspectives, showcased breakthroughs in technical routes including superconductivity, photonic quantum, and ion trap, and shared a large number of specific industrial scenarios, industrial system construction, and prospects of heterogeneous computing. The participants jointly explored the future of technology industrial investment.
The following dialogue is organized and edited by 36Kr:
Jin Tao | Partner, Gaohe Capital (Host)
Zhao You | Executive President, Neolithic Autonomous Vehicles
LIN Tianwei | Algorithm Scientist, Mochi Intelligence
Jin Tao: Hello everyone! I'm Jin Tao, Partner of Gaohe Capital. Gaohe is a boutique investment bank focusing on the new economy sector. Over the past year, we have completed more than 80 billion US dollars of financing projects in the fields of AI, embodied intelligence and cutting-edge technology, including outstanding enterprises such as MiniMax, Sudu Technology, Neolithic Autonomous Vehicles, Mochi Intelligence, Tianji Intelligence, Volant, Buchou Quantum, Inverse Matrix Technology and others. The two guests present today, Zhao You, Executive President of Neolithic Autonomous Vehicles, and LIN Tianwei, Chief Algorithm Scientist of Mochi Intelligence, are both from partner enterprises that Gaohe has served deeply for a long time.
At the 36Kr conference last year, I also discussed the commercial implementation of autonomous driving with Zhao You (Will). At that time, Neolithic had just crossed the threshold of 10,000 units in operation; a year later, the operating scale of its autonomous vehicles has exceeded 20,000 units and is still growing continuously. This is a very representative industrial sample for exploring how AI can truly enter the real world.
Tianwei currently serves as Chief Algorithm Scientist of Mochi Intelligence. He used to be the head of end-to-end pre-research and embodied operation at Horizon Robotics, and led the R&D of Horizon's first version of end-to-end real vehicle solution. Now he continues to focus on embodied operation and robot intelligence.
It is quite interesting that both of you have deep connections with the automotive and autonomous driving sectors, and now stand at different positions where AI enters the real world: one has promoted autonomous driving to real operation on a scale of tens of thousands of units, and the other has moved from end-to-end autonomous driving to embodied operation. Taking this opportunity, we want to discuss a more specific question: when AI leaves the screen and enters the complex physical world such as roads, factories, hotels and homes, do we only need stronger models, or a complete set of system capabilities?
The first question I would like to ask both of you. Autonomous driving can be said to be one of the earliest, largest-scale and highest-investment long-term experiments for AI to enter the physical world. Looking back at the development over the past ten years, what are the most important inspirations it brings to this round of Physical AI and embodied intelligence?
Zhao You: I totally agree! Autonomous driving is indeed the largest-scale and highest-investment experiment for AI to enter the physical world, and startups and investment institutions have invested a lot of capital. Combined with our practice, I have several insights:
1. After AI enters the physical world, the first constraint comes from the ontology. What it faces is not unlimited computing power, but the need to deploy models with limited parameter scales to the vehicle end for operation. The ontology itself is constrained by customer costs and product requirements. In the final analysis, enterprises must truly understand customers and operations, define the ontology properly, and form a data flywheel on this basis, so that the system can continue to operate and iterate. This is a fundamental insight formed from our years of practice, and it can also be said to have gradually become the industry consensus.
LIN Tianwei: From the perspective of physical AI, intelligent driving and embodied intelligence have strong similarities. Looking back on the development of intelligent driving over the past 10 years, there is an important enlightenment for embodied intelligence: the evaluation of effect should not only focus on the optimal performance in a certain scenario, but also pay long-term attention to the stability of the system's continuous operation in general scenarios and the ability to handle conventional problems.
As early as 10 years ago, we could see some demo vehicles, for example, completing autonomous straight driving, turning and driving on specific routes in North America. But until the last one or two years, we have gradually seen vehicles achieve hundreds of thousands of hours of non-intervention driving in a truly unmanned state. This process is mostly about solving the last 1% of problems. It is not difficult for vehicles to have basic driving capabilities. The real difficulty lies in long-term and stable continuous operation, building a data flywheel and solving the closed-loop problem under abnormal conditions.
Compared with the model architecture changes that are often discussed, this is more like a long-term closed-loop project, which may not be that eye-catching, but it is an indispensable part if we want to achieve large-scale applications.
From the enlightenment of intelligent driving to embodied intelligence, we need to pay more attention to the final evaluation indicators and actual effects, such as long-term stability. Intelligent driving pays attention to the intervention rate, and embodied intelligence may also form similar indicators in the future, such as what tasks can be completed, how long it can run without manual intervention, etc. On this basis, we can further discuss specific model solutions.
Intelligent driving has experienced the evolution from traditional rules, modular AI to end-to-end models, and new technical paradigms may emerge in the future. From the perspective of the indicators that need closed-loop and the acceptance methods, the underlying logic has not changed. What is more important is to clarify what the evaluation criteria are from the very beginning.
LIN Tianwei: I fully agree with the evolution direction from rules to models. After truly entering the era of massive data, we can clearly see the emergence of capabilities in recent years, and we can also see that the scaling law plays a role in real scenarios, which is a very exciting thing for us.
Before the rapid development of large models and autonomous driving, people did not have such strong confidence in this matter; now we have clearly seen the effect brought by large-scale data and models, and I think this has strong reference significance for AI to enter the physical world in the future.
Jin Tao: I would like to ask Will another question. Last year, Neolithic just crossed the 10,000-unit mark, and now it has exceeded 20,000 units. As the operating scale continues to expand, what new insights and discoveries have you gained?
Zhao You: I think there are several different points.
First, when the scale was at the 1,000-unit level, the proportion of R&D investment was very high, and the cost per unit spread to each shipment was also very high. Therefore, the company paid great attention to ROI in R&D management and tried to control the total investment. When the scale reaches tens of thousands of units, we find that R&D investment is no longer the largest part of the cost. The daily operation and maintenance, power consumption, vehicle depreciation of each vehicle, as well as the costs caused by algorithm-related intervention and maintenance have become more prominent. These costs will in turn affect the R&D direction, making R&D focus more on reducing the operating cost per unit. This is a very obvious change after mass production and commercialization.
Second, related to this, when the scale was at the 1,000-unit level, we paid more attention to R&D capabilities; when the scale becomes larger, organizational capabilities will become a greater challenge. After the scale of the fleet or ontology expands, the corresponding number of personnel increases, which brings about management and coordination problems. The capability building of the company will further evolve from a pure R&D issue to an organizational issue. This is a very obvious change in the process from R&D to mass production.
Finally, the expansion from 1,000 units to 10,000 units puts forward higher requirements for generalization capabilities. As we enter more and more cities and scenarios, long-tail problems gradually change from occasional situations to daily problems.
Therefore, the requirements for data flywheel and model capabilities are also getting higher and higher, which is a point we have deeply experienced in the past year.
Jin Tao: Tianwei, you used to be in charge of end-to-end pre-research and embodied operation at Horizon Robotics, and also led the R&D of the first version of end-to-end real vehicle solution. Autonomous driving has formed a relatively mature technical closed-loop and verification system over the past ten years. When these experiences are transferred to embodied intelligence, what can be reused and what are the essential differences?
LIN Tianwei: First of all, intelligent driving and embodied intelligence are two very different things. The hardware faced by autonomous driving is relatively mature, including the vehicle itself, the installation method of on-board sensors, etc., which have formed a relatively stable form at present; while the hardware form of embodied intelligence is still in a relatively non-convergent stage.
Secondly, there are also significant differences in task definition. Autonomous driving is relatively standardized, and the task goal is very clear: to operate together with other traffic participants in the road network and reach the destination safely. The task itself is relatively convergent, and one of the core principles is to avoid collision with other objects, with safety as the priority.
Embodied intelligence is just the opposite in these two aspects. We hope that embodied intelligence can solve general tasks, and the potential scope of such tasks is very large, so it is difficult to define them in a highly structured way like intelligent driving. For example, in intelligent driving, we can distinguish static tasks and dynamic tasks, but in embodied intelligence, it is difficult to complete the general definition in the same way.
In addition, embodied intelligence has the feature of strong physical contact. We need to operate objects instead of avoiding them. In the operation process, we will encounter a large number of problems related to physical properties, such as fluids, flexible objects, etc., and the differences will be very obvious.
For example, two pieces of clothing may look the same, but due to different surface smoothness or weight, if the strategy does not have sufficient adaptability, it may not be able to complete the task when facing a different object. I think this is a very important difference between the two types of systems.
Jin Tao: Tianwei, recently Mochi Intelligence released achievements such as MORPHI KINO and MoRA. One important change is that robots are moving from completing a single atomic action to more complex long-horizon tasks. But what key thresholds need to be crossed between "being able to complete long-horizon tasks" and "being truly usable and capable of continuously completing real work"?
LIN Tianwei: Single-point tasks and long-horizon tasks are significantly different. Taking a single task such as folding clothes as an example, we usually learn the task process through imitation learning, so that the model knows what actions to perform at different stages of the task.
But this approach will bring many problems. In some cases, the model does not know whether the task has been completed, because it only imitates actions and does not have the ability to judge the completion state of the task.
If the model only fits for a single-point task, its ability to follow task instructions may be insufficient, and over-fitting may also occur. Long-horizon tasks usually adopt the agent architecture, which is equivalent to a "brain" at the System 2 layer that breaks down vague instructions into more specific task instructions.
Then the embodied model at the System 1 layer is responsible for action execution. The main bottleneck now is not the task decomposition of the "brain", but the execution layer: the model's ability to follow instructions and generalization are still not good enough.
For example, some models have been trained to grasp apples, but they may not succeed when grasping other fruits in a different scenario, which shows that the generalization ability is still insufficient. If the action model of System 1 cannot follow instructions stably, the brain cannot schedule effectively. Therefore, the first biggest challenge is to further improve the Action model of System 1 in terms of instruction generalization, object generalization and instruction following.
In addition, long-horizon tasks are composed of multiple sub-tasks in series, so it is necessary to continuously monitor the progress of each task, including whether it is completed, whether an error occurs, etc. On this basis, we build the harness agent related modules to give feedback and make adjustments according to the execution and scheduling situation.
For example, when an error occurs, the task may be terminated directly after the single-point model fails; but long-horizon tasks will definitely encounter failures in intermediate links. If a task has ten links, it is impossible for each link to reach a 100% success rate. Therefore, the system must be able to adjust the strategy when a failure occurs, re-plan the instructions and continue to complete the task. This is also one of the key research directions for us in the future.
Jin Tao: Will, Neolithic has moved from technical verification to real operation on a scale of tens of thousands of units. After completing a large number of real tasks every day, the criteria for evaluating the system have further shifted from "whether it can run" to reliability, safety and operation efficiency. How do these requirements in turn affect your technical route and system architecture?
Zhao You: We also found that the technical architecture has changed significantly compared with the early stage. Reliability is directly related to safety in autonomous driving, so it is the top priority indicator at the end side. The end-side algorithm must first ensure the safety of intelligent driving performance, but logistics autonomous driving will encounter a large number of corner cases.
As Tianwei said just now, road driving is relatively standardized and the rules are clearer; but for logistics, the scenarios at both ends are more complex, and it may be necessary to enter a large number of non-standard end scenarios such as residential areas, wholesale markets, warehouses, and factories. In these scenarios, the vehicle speed is low, the form of safety risks will change, but the environmental complexity is higher, and the requirements for intelligent capabilities are also higher, which sometimes even exceeds the range that the end-side computing power can handle independently.
Therefore, we have added a cloud-side model in addition to the end-side model, and the cloud-side model and the end-side model jointly handle more complex scenarios; at the same time, there are remote intervention personnel as the third layer of guarantee. Such an operation-oriented technical architecture supports the current scale of 10,000 units, and prepares for the operation of hundreds of thousands of vehicles across the country and even around the world in the future.
In short, reliability and safety are the most important, so key capabilities must be deployed at the end side to ensure safety and reliability in a low-latency manner. Facing more complex scenarios, we further improve system performance through cloud and remote support.
Jin Tao: In the past, we were used to the path that technology matures first and then lands in commercialization; but in many technology fields today, technology and implementation often promote industrial development through mutual evolution.
I would like to ask both of you, in your respective fields, how do you realize the co-evolution of technology and scenarios?
LIN Tianwei: I think the technical development of embodied intelligence and scenario entry cannot be independent of each other. It is not an ideal path to spend two years polishing the technology to maturity first and then look for scenarios. The characteristic of embodied intelligence is that if you do not enter real scenarios, you may not even know what tasks really need to be completed. For example, when we enter the hotel scenario now, we find many real demands. Take wiping tables as an example: in the laboratory, it may just be a simple wiping action, but hotel cleaners have a clear set of processes and standards.
After entering the scenario, on the one hand, we can clarify what the SOP of these tasks is and what standards need to be met; on the other hand, we can collect data in real scenarios, form a closed loop, and iterate the model effect more directly.
In the long run, such real scenarios can also help us improve the model's generalization ability to complex environments. For example, if we want to achieve good performance in the home scenario in the end, hotels can be used as a good intermediate verification scenario: the space layout has certain similarities with homes, but it is easier to enter in commercial terms. Directly entering homes will also involve issues such as privacy and management, so we will first complete the closed loop in the hotel scenario and continuously polish the overall solution.
Zhao You: Neolithic is quite special in the combination of technology and scenarios. We have been rooted in logistics from the very beginning and chosen logistics as our core scenario. I have two insights here.
1. Logistics and