Palms, desktops and wheels: On-device AI kicks off the "battle for entry points"
When the industry is still caught up in endless debates over the parameter scale and training computing power of cloud large models, a far more decisive war has quietly broken out on the terminal devices closest to users.
From car manufacturers intensively releasing intelligent driving solutions equipped with edge-side large models, to mobile phone manufacturers embedding 100-billion-parameter models into flagship chips, and then to AI PCs and smart wearable devices all putting up the banner of "local native intelligence", all signals point to the same conclusion: The main line of AI competition is shifting from the "cloud computing power race" to the "scramble for edge-side entry points".
The development history of the Internet is essentially a history of entry point iteration. Whoever masters the first entry point for users to access the digital world will take the dominant control over traffic, data and business. In the PC era, the dominant entry points were browsers and search engines, while in the mobile Internet era, they were super Apps and app stores.
Entering the AI era, intelligence is no longer limited to screens and clicks, but penetrates into every physical scenario such as driving, perception and interaction, and the definition of entry points is being completely rewritten. This time, the deciding factor of the war does not lie in remote data centers, but in every terminal device within easy reach.
The cycle of entry points makes the edge side a must-contested territory in the AI era
Every paradigm shift of technology is accompanied by a reconstruction of entry points.
In the 1990s, when the Internet just entered the public view, web portals were the absolute traffic center, and users accessed the entire network through the homepages of Yahoo and Sina; after the rise of search engines, Google and Baidu became the core entry points for information distribution, and people obtained services through keywords; when the mobile Internet wave hit, super Apps replaced websites, WeChat and Alipay became the main channels for people to access digital life, and the traffic distribution right shifted from browsers to the application ecosystem.
The replacement of each generation of entry points follows the same underlying logic: whoever is closer to users, responds faster and delivers a more seamless experience will become the first choice of users.
At the beginning of the arrival of the AI era, the industry generally believed that cloud large models would become the new generation of entry points. Users send instructions to the cloud large model through natural language, and the large model completes intention understanding, task disassembly and service scheduling. However, with the deepening of scenario implementation, the inherent shortcomings of the cloud mode have become increasingly prominent, and more and more people realize that pure cloud AI cannot support real-time intelligence in the physical world.
The first unavoidable problem is latency. Cloud reasoning relies on network transmission. Even under 5G networks, there is a delay of tens to hundreds of milliseconds from data upload to result return, and the delay will be higher when the network fluctuates. For scenarios with extremely high latency requirements such as autonomous driving, industrial quality inspection and real-time AR, a delay of hundreds of milliseconds may lead to safety accidents. When an autonomous vehicle is driving at high speed, obstacle recognition and path planning require millisecond-level response. Once network jitter causes delay of cloud instructions, the consequences will be unimaginable.
The second core pain point is privacy and data security. Under the cloud mode, users' voice, image, location and behavior data need to be continuously uploaded to the server for processing. The on-board camera of smart cars will shoot the road environment and the images of drivers and passengers, and the smart assistant of mobile phones will read users' address books, schedules and albums. These data contain highly sensitive personal information. Frequent data transmission not only increases the risk of leakage, but also exposes enterprises to increasingly strict global data compliance supervision. The requirements for automotive data security management require that important data shall not leave the country and sensitive data shall be processed locally, which is inherently contrary to the pure cloud architecture.
The third practical constraint is cost. The computing power cost of large model reasoning is extremely high. When the scale of terminal devices reaches hundreds of millions levels, calling cloud computing power for every interaction will generate unbearable bandwidth and computing power expenses. For car manufacturers, if the intelligent driving perception of millions of vehicles relies on cloud processing, the cost of traffic and servers alone will devour a large amount of profits.
The emergence of edge-side AI is precisely a systematic solution to these pain points.
The so-called edge-side AI refers to deploying the lightweight large model directly on the local terminal device, all reasoning calculations are completed inside the device, and the original data does not need to be uploaded to the cloud.
It brings three core values that the cloud cannot replace:
First, extremely low latency. Data does not need to travel back and forth to the cloud, and the reasoning process is completed locally, so the response speed can be compressed to the millisecond level, which fully meets the rigid requirements of scenarios such as autonomous driving and real-time interaction.
Second, native privacy. All original perception data and personal behavior data are processed locally, and only desensitized incremental information is uploaded when necessary, which reduces the risk of data leakage from the source and is more in line with the compliance principle of data minimization.
Third, available offline. Smart functions can still operate normally in underground garages, tunnels and remote road sections without network signals, and will not fail due to network disconnection.
These three characteristics determine that edge-side AI is by no means a supplementary solution to the cloud, but the only way for AI to move towards the physical world and integrate into real scenarios. For this reason, it has become the core battlefield of the new generation of entry point wars — when intelligence becomes ubiquitous and omnipresent, the terminal closest to users naturally becomes the first entry point to undertake user intentions and distribute intelligent services.
Three major tracks are being raced for, and the war spreads from the palm of the hand to the wheels
The war of edge-side AI has broken out simultaneously in the three major tracks of smart vehicles, consumer electronics and general IoT. Among them, smart vehicles have become the number one "testing ground" because of their most complex scenarios, highest value and most rigid demand for edge-side capabilities.
Cars are evolving from transportation tools into mobile smart spaces. And the edge-side large model is exactly the core technical base supporting this transformation.
In the field of intelligent driving, the value of edge-side AI has reached a consensus in the industry. In the past, assisted driving systems relied on rule-based algorithms and cloud assistance, and often failed to cope with complex urban scenarios, special-shaped obstacles and extreme weather. The multi-modal large model deployed on the edge side can complete the full-link closed loop of environment perception, target recognition, behavior prediction and path planning locally, truly realizing "data does not leave the car, decision-making does not wait".
The current mainstream technical trend in the industry is to adopt the "fast and slow dual thinking" architecture. The lightweight end-to-end model is responsible for basic controls such as millisecond-level emergency braking and lane keeping to ensure driving safety; the larger-scale edge-side large model is responsible for semantic understanding and decision reasoning for long-tail scenarios such as complex intersections, unprotected left turns and construction sections. At the same time, the on-board world model is beginning to be implemented, and vehicles can locally deduce the road condition changes in the next few seconds and make decisions in advance, completely getting rid of the real-time dependence on cloud computing power.
In June 2026, Li Auto released Mach Mind-Edge. Within the language intelligence system, Mach Mind-Edge works in coordination with the cloud model Mach Mind-Pro. Mach Mind-Pro is responsible for general reasoning and complex task scheduling, while Mach Mind-Edge as an edge-side model is mainly responsible for local tasks such as real-time human-vehicle interaction and autonomous vehicle control.
The 7B edge-side model of AutoX uses optimized architectures such as MoE sparse activation to control the reasoning delay at about 100 milliseconds, covering professional knowledge such as road topology, traffic rules and vehicle behavior, while avoiding redundant calculations. When a sudden road accident occurs, the system can complete over-the-horizon perception, calculation of the affected range and optimal detour path planning within a few seconds.
What has more entry point imagination than intelligent driving is the whole-vehicle intelligence brought by the integration of cabin and driving. As the central computing platform replaces distributed ECUs, a single high-computing-power SoC can carry both intelligent driving decision-making and cockpit interaction. The edge-side large model has become the intelligent hub of the whole vehicle, perceiving road conditions and planning routes externally, and internally understanding the voice, gestures and even emotions of drivers and passengers, and scheduling all services in the vehicle.
This means that the entry point value of cars is no longer limited to navigation and entertainment, but becomes a mobile entry point for users' travel, life and work. The premise of all this is that the intelligence capability must be localized — no one wants to be unable to use the voice assistant due to network disconnection in a tunnel, let alone have their in-vehicle conversations and travel data continuously uploaded.
Data security and compliance are even more core plus points for edge-side solutions. Driving images, driving and riding audio and video, and vehicle trajectories are all sensitive data. Local processing on the edge side reduces external data transmission from the source, greatly reducing the risk of data leakage and compliance risks. This is also the fundamental reason why more and more car manufacturers take "data does not leave the car" as their core selling point.
If cars are the high-value battlefield of edge-side AI, then mobile phones and AI PCs are the largest-scale mass battlefield.
Smartphones are already the undisputed entry point of the mobile Internet, and edge-side large models are upgrading this entry point from "application distribution" to "intention distribution". In the past, users needed to open different Apps to complete different tasks. In the future, users only need to send instructions through the local large model, and the model will directly schedule system capabilities and application services.
The miniaturization of edge-side models is gradually becoming an industry consensus. Apple's Apple Intelligence also adopts a 3B parameter design, which can complete tasks such as text summarization, information extraction and cross-application operation on the device. Zhou Wei, Vice President of vivo and Dean of vivo AI Global Research Institute, put forward a new development direction — giving up the pursuit of "general artificial intelligence" and launching a lighter 3B edge-side multi-modal reasoning large model. This model only occupies 2GB of mobile phone memory with power consumption of about 750 mA, and its reasoning capability is comparable to that of 7B or even 10B models.
The rise of AI PCs further extends the competition for edge-side entry points to productivity scenarios. PC chips equipped with independent NPU can run large models locally, realizing functions such as offline document summarization, table analysis, code generation and meeting minutes. For business users, local processing means that commercial confidential data does not need to be uploaded to the cloud, and security is greatly improved. When AI capability becomes the standard configuration of PCs, operating system and chip manufacturers will master the intelligent entry point right in productivity scenarios.
A large number of IoT terminals such as smart watches, smart homes, security cameras and industrial edge devices outside mobile phones and cars are forming the long-tail battlefield of edge-side AI.
The common characteristics of these devices are limited computing power, unstable network environment, and extremely high requirements for privacy and power consumption. The lightweight small edge-side model can complete basic intelligent functions such as face detection, anomaly recognition and voice wake-up at extremely low power consumption without relying on the cloud. Although the value of a single device is not high, the aggregation of massive devices forms a huge intelligent entry point network covering families, industries and cities.
Edge-side AI is not as simple as "stuffing the model into the terminal"
Under the hot track, the industry has gradually reached a consensus: the threshold of edge-side AI is much higher than imagined. It is not simply compressing the cloud model and putting it into the chip, but a full-system tough battle involving algorithms, software, hardware, engineering and security.
The first thing edge-side AI has to face is the "impossible triangle" composed of computing power, power consumption and cost.
The physical constraints of terminal devices are extremely strict. Vehicle chips need to work stably in a wide temperature environment from -40℃ to 85℃, and their power consumption is strictly limited by the vehicle power supply and heat dissipation system; mobile phone chips need to balance battery life and hand feel, with extremely limited heat dissipation space; IoT devices are even powered by batteries, with a power consumption budget of only milliwatt level.
Under such constraints, it is impossible to stack computing power without limit. However, large models, especially the Transformer architecture, have extremely huge demand for computing power and memory bandwidth. The massive memory access operations brought by the self-attention mechanism often make the chip have computing power but cannot feed enough data, forming the "memory wall" bottleneck.
The industry's solution is model lightweighting: INT4/FP8 low-bit quantization, structured pruning, knowledge distillation, sparsification, MoE Mixture of Experts... a series of technical means are used to compress the model volume, so that the large model can run under limited computing power. But compression comes at a price. Excessive quantization and pruning will lead to the decline of model accuracy, resulting in recognition errors and decision deviations in long-tail scenarios.
This is particularly fatal for autonomous driving. The model performs well under conventional road conditions, but once encountering special-shaped obstacles, extreme weather and rare traffic scenarios, the recognition accuracy of the lightweight model may drop significantly. How to maintain the accuracy of key scenarios within the limited computing power and power consumption budget is the core exam question that all edge-side AI teams have to face.
If model lightweighting is the basic question, then software-hardware collaborative optimization is the advanced question that widens the gap.
Many teams will encounter an awkward problem: chips with very high theoretical computing power have very low effective computing power when actually running large models. The reason is very simple: large model operators are complex and memory access intensive. If you just run with a general reasoning framework, you cannot give full play to the real performance of the chip.
To maximize the efficiency of edge-side large models, you must start from the bottom layer. Customize and develop operators for the chip architecture, optimize memory scheduling, connect the reasoning framework with the underlying driver, and complete operator fusion and graph optimization. This requires deep cooperation among the algorithm team, compiler team, chip team and driver team, and consider hardware adaptation from the model structure design stage, rather than moving the model to the chip after training.
This puts forward new capability requirements for enterprises. In the past, the algorithm team was only responsible for training models, and the chip team was only responsible for making hardware, and the two were relatively independent. But in the edge-side AI era, the two must be tightly coupled. Car manufacturers cannot only do algorithm integration, they must go deep into the chip operator and scheduling level; chip manufacturers cannot only sell hardware, they must provide a complete software stack and model adaptation tools; large model companies cannot only output cloud APIs, they must have the engineering capability of edge-side deployment and quantization optimization.
The deeper challenge behind it is talents and organization. Edge-side AI requires compound talents who understand large model algorithms, compilers, chip architectures and automotive-grade engineering at the same time, and such talents are extremely scarce in the industry. Organizationally, traditional department walls will seriously hinder software-hardware collaboration, and enterprises must establish a cross-departmental joint iteration mechanism to shorten the tuning cycle.
The implementation of edge-side AI, especially in the field of autonomous driving, security and compliance are two major hurdles that cannot be bypassed.
The first is functional safety. The automotive industry has strict ISO 26262 functional safety standards, and the highest level ASIL-D requires extremely low system failure probability. But large models are essentially probabilistic outputs with black-box features. The decision-making process is difficult to fully explain, and it is difficult to exhaust all failure boundaries. This makes it difficult for edge-side large models to directly pass the highest level of automotive safety certification.
The current response idea of the industry is "large model + safety redundancy". The large model is responsible for complex decision-making, while retaining the traditional rule-based algorithm as a safety fallback. Once the output of the large model is abnormal, the safety system will take over immediately. But how to define the power and responsibility boundary between the two, and how to verify the safety of large models in all extreme scenarios, there is no mature industry standard so far, and it is still in the exploratory stage.
The second is data security. Many people think that local data processing is absolutely safe, which is actually a misunderstanding