HomeArticle

The competition in the second half of the AI industry is not about who is better at chatting.

晓曦2026-09-10 15:10
The real world is the ultimate exam paper.

Text by Wang Yi

Over the past two years, AI has drastically reduced the cost of information generation, but the credibility of the answers has also been declining rapidly. Travel guides may recommend scenic spots that have been closed down, and the local must-try restaurants may not live up to their reputation. Many seemingly complete travel plans can only be used as references by users, who dare not follow them directly for execution.

As answers become increasingly cheap, whether they are authentic has become a new problem.

The essence of the problem is not that large models are not smart enough, but that most AIs lack the ability to understand the real world. Information can be retrieved and generated, but the real world is not only composed of texts. It is difficult for AI to judge the actual queuing situation of a restaurant, the experience difference of a block in different time periods, and it is also difficult to understand important variables affecting decision-making such as sense of trust, spatial atmosphere, and individual preferences.

Such capabilities are precisely built on the long-accumulated spatiotemporal data and map infrastructure of AutoNavi. Traditional maps solve the problem of how locations are connected, and navigation solves the problem of how to arrive. Spatial intelligence further begins to describe the real world itself: 3D spatial structure, real-time changing pedestrian and vehicle flow, and the dynamic relationship between people and the environment.

Therefore, AutoNavi today can not only answer "what the world looks like", but also understand "what state it is in at the moment, how it will change next, and what actions should be taken". For AutoNavi, spatial intelligence is not just a new technology added to maps and navigation, but also a change of capability paradigm, which expands the boundary of AutoNavi's services for the real world.

01. Real World Understanding Capability, the New Barrier for AI Applications

Compared with other AIs that rely on text reasoning, AutoNavi tries to implement spatial intelligence into specific travel scenarios.

For users, a travel plan with practical reference value requires sufficiently credible information, recommendations that fit individual aesthetic preferences, the ability to restore on-site experience in advance, and at the same time take into account complex realistic variables such as time, weather, travel companions, and traffic conditions.

Therefore, a list that fits the real world is not a simple sorting of locations, but a simulation of the real world. Driven by spatial intelligence, AutoNavi Street Exploration List 2026 has also achieved a comprehensive upgrade around world simulation.

Previously, the internet review ecosystem has long been interfered by noises such as marketing traffic. AutoNavi Street Exploration List uses the real navigation store arrival behavior that cannot be forged to form users' most direct judgment on a store, which gives the list a credible basis. On this basis, the algorithm of AutoNavi Street Exploration is further subdivided to identify whether each store visit is a special trip or a passing-by, a returning customer or a new customer, a local resident or a tourist.

The logic is obvious: traffic is not the goal, and the gold content of traffic is what matters. A store that is worth customers traveling several kilometers to visit specially is certainly not equivalent to a store that is just passed by when it comes to "being worth recommending".

In addition, behavior data can prove real store visits, but people who arrive at the store also have differences. If everyone has the same scoring weight, some professional judgments will easily be submerged by the average value.

The solution of the Street Exploration List is to introduce professional scoring. For lists with strong professionalism such as coffee shop lists, the votes of professionals with their actual visits will get higher scoring weight. In today's era where AI recommendations are generally homogeneous, only by giving subjective variables like taste higher weight can we pick out those unique and high-quality venues from the recommendations that cater to the average public taste.

The list can tell users where to go, but cannot answer what a store actually looks like. In the past, users could only imagine the scene through text and image reviews, and the information gap between online and offline became the psychological gap after users actually arrived at the store.

Flight Street View 2.0 tries to eliminate this information gap. Relying on AutoNavi ABot-Earth 0.7 world model, it turns the destination into a 3D space that can be freely explored, allowing users to enter the real space before departure, freely adjust the viewing angle and position, and get an immersive on-site feeling on the screen.

This experience of viewing the scene first and then setting restores the decision-making from abstract scores to concrete scenarios, responding to users' anxiety about the unknown.

In addition to Flight Street View 2.0, the new mine avoidance guide launched by AutoNavi Street Exploration List also uses deterministic spatiotemporal calculation to fight against the uncontrollable uncertainty in travel. The mine avoidance guide will combine real-time dynamic perception and spatiotemporal deduction to verify the whole trip, identify risks such as time conflicts, congestion and store closure in advance from 15 dimensions including business hours, trip rhythm, parking prompts and crowd prediction, and inform users of the pits they may step on and executable suggestions in advance.

After actually setting off and arriving at the destination, Navigation Live can not only understand users' language, but also perceive users' position, direction and surrounding environment in real time, and understand the street view scene through the camera to give suggestions highly related to the current situation.

The essence of Navigation Live is an embodied intelligent agent with "spatiotemporal context + on-site perception + end-to-end action capability". For example, when you drive or walk in a complex old block in an unfamiliar city, there are three extremely close narrow alley entrances appearing at the same time ahead. Traditional navigation will only prompt "turn right 50 meters ahead", and you will most likely go wrong; while Navigation Live can directly understand the scene: "Please drive into the alley on the right side of the convenience store with red lanterns hanging, and avoid the van that is unloading goods on the left".

02. Spatial Intelligence, the Base to Understand the Real World

From the real consumption decision-making provided by the Street Exploration List, to Flight Street View allowing users to see the destination clearly in advance, to the mine avoidance guide that deduces the trip in advance and Navigation Live that accompanies users into the real world, AutoNavi extends its understanding of the real world from a single list to the complete full-day travel link.

Behind this lies AutoNavi's capabilities of 3D spatial representation of spatial intelligence, real-time dynamic perception and spatiotemporal deduction.

First of all, the starting point of all cognition is to build a 3D space that can be understood by machines. Traditional navigation applications are used to marking locations with points, labels, text and image information. But the real world, such as the width of streets, the orientation of stores, and walking accessibility, is built on spatial relationships. AutoNavi completely digitizes these physical attributes through the 3D native city world model ABot-Earth 0.7, tells users the spatial feeling when they are in it, and lays a foundation for subsequent simulation and prediction.

After having the static 3D world, various sudden changes in the real world, such as temporary road construction and temporary store closure, will not be written into the database in advance. Dynamic perception is equivalent to the sensory system of this world model, which continuously receives and feeds back the latest signals of the real world.

In addition, the real world is always in flow, and the pedestrian flow and vehicle flow form a huge dynamic fluid system. Spatiotemporal deduction allows AI to see how the world is operating.

Navigation Live is a typical product that demonstrates the collaborative work of the three capabilities mentioned above.

Imagine a daily scene of driving through the city: during the driving process, the end side continuously scans the surrounding environment with millisecond-level dynamic perception to complete the real-time reading of the real world, which can not only identify the flashing brake lights ahead and the roadblocks set on the road, but also capture the newly opened cafe on the street, and completely transmit "what is happening right in front of us" to the system.

The spatiotemporal deduction on the cloud undertakes the signals collected in real time to make beyond-visual-range global predictions. For example, predicting whether oncoming vehicles will come out of the curve blind area, evaluating how much obstruction a road occupation construction will bring to the traffic flow in the next few kilometers, and revealing potential risks that have not entered the field of vision in advance.

The 3D spatial representation undertakes the anchoring and integration between digital information and physical scenes. In real scenarios, the vehicle keeps bumping and the perspective of the device is also moving continuously. Relying on the precipitated 3D spatial base, Navigation Live can make all kinds of prompt information firmly anchored on the real street view, so that the picture does not drift and the information does not shake, and seamlessly superimpose route prompts, risk warnings and surrounding location introductions into the real field of vision.

03. AutoNavi, the Undervalued Player in the AI Era

Since the beginning of this year, industry forums and research institutions have pointed out that the AI industry has gone through the stage of competing for parameters and text generation effects. As the conversational capability of large models becomes increasingly mature, the new proposition of the industry has transformed into: AI can no longer only live in the dialog box. To truly release its value, it must step into the physical world.

Fei-Fei Li, known as the "Godmother of AI", once proposed in the early stage that large models have profound linguistic knowledge but lack personal perception of the physical world. In her latest interview, she mentioned again: In my opinion, without spatial intelligence, general artificial intelligence is incomplete.

The real capital investment from the industrial side also confirms the above judgment. World Labs founded by Fei-Fei Li has received an investment of 230 million US dollars, focusing on developing AI with "spatial intelligence"; Yann LeCun, former chief AI scientist of Meta, left Meta last year and founded AMI Labs, raising more than 1 billion US dollars to build a world model system; Wang Xingxing, founder of Unitree Robotics, also stated at this year's World Robot Conference that the biggest bottleneck of embodied intelligence is insufficient generalization capability at present, for which the company is focusing on promoting the research and development of physical AI robot models to accelerate the completion of the shortcomings of the embodied "brain".

Automobile manufacturers begin to talk about physical AI. Intelligent driving no longer only pursues the smoothness of cockpit conversation, but focuses on tackling world models to enable vehicles to understand road space; practitioners of embodied intelligence also realize that the bottleneck of robots is not all in dialogue interaction. Understanding 3D environments, grasping spatial relationships and predicting object movements are the unavoidable links in the implementation process.

This shows that whether it is autonomous driving, humanoid robots, urban digital governance or local life services, all AI applications that need to connect with the real physical world cannot do without the support of spatial intelligence.

Large models can generate countless beautiful answers in batches, but the real world is the ultimate test paper for testing the value of AI. In the past, we judged AI by whether it can output wonderful content fluently; in the future, the key to distinguish the capability of AI is to see to what extent it can understand the constantly changing real world.

It is under such an industry background that the value of AutoNavi has been highlighted again. Understanding the real world does not mean knowing more location information, but being able to judge what a specific person may see, whether it is crowded, how long he will stay, and where he will most naturally go next, under the specific conditions of person, time, location, travel companions, weather and preferences. What it ultimately needs to answer is not only "where it is and how to get there", but also "what will happen", "what the experience is like", "who it is suitable for" and "why".

Without the accumulation of spatiotemporal data such as 2D networks, 3D spaces, time changes, pedestrian flow, vehicle flow and navigation behaviors, it is difficult for AI to understand the real world, and these are exactly the paths AutoNavi has been walking on in the past. The "sense of language" of space and time precipitated by AutoNavi over time cannot be quickly achieved by money or traffic, and it is also the real moat of AutoNavi.