HomeArticle

Waymo effectively pronounces the death sentence on the pure vision technical route: Tesla can implement FSD, but faces enormous difficulties in developing driverless Robotaxi.

智能车参考2026-08-07 17:02
L4 requires the last few "9"s

"Pure vision cannot achieve full autonomous driving."

In his latest public speech, the Co-CEO of Waymo has "sentenced pure vision to death".

He admitted that human drivers can drive only relying on their eyes, which proves that the camera-based solution is feasible. Pure vision systems can indeed develop rapidly, and even quickly achieve impressive driving performance.

But the problem is that the goal to be achieved by L4 autonomous driving is not just driving like a human, but to reach a safety level far exceeding that of human driving.

In his view, the biggest challenge of the pure vision route is that performance improvement may slow down earlier and reach its peak prematurely before reaching this goal.

He did not mention the word "Tesla" at all, but every sentence was directed at Tesla.

Pure vision can reach 99%, and then what?

This speech was delivered at YC Startup School 2026, a free online course and community project hosted by Y Combinator, the top Silicon Valley startup accelerator.

Co-CEO of Waymo — Dmitri Dolgov, gave a nearly 50-minute speech titled "The Demo Is Only 1% Of The Work", systematically explaining why autonomous driving has not truly entered the driverless era for a long time.

This title has actually pointed out the core judgment of the entire sharing: the hardest part of autonomous driving is to make the system operate continuously and stably on millions and tens of millions of miles of real roads.

Dolgov shared an experience from the early days of Waymo.

Around 2010, the Google autonomous driving project, the predecessor of Waymo, completed a large number of tests in a relatively short period of time, including 100,000 miles of autonomous driving and multiple long-distance routes with zero human intervention.

From the perspective of technical demonstration, autonomous driving seems to have made a breakthrough. But after that, Waymo spent more than ten years before officially operating Robotaxi for the public.

He believes that the reason is that the real difficulty of autonomous driving lies in the improvement of reliability. For the convenience of understanding, Dolgov put forward the concept of "nines":

It requires a lot of engineering investment to improve the reliability of the system from 90% to 99%; when continuing to improve from 99%, every additional "9" after the decimal point will increase the difficulty rapidly.

Autonomous driving faces the real world. After the vehicles run hundreds of thousands or even millions of miles per week, problems that originally occurred with extremely low probability will gradually become problems that must be faced in the operation process.

Therefore, autonomous driving cannot only focus on the "average performance". What really determines whether the driver can leave the driving position are those extreme scenarios — which is also the starting point of Dolgov's subsequent discussion on the pure vision route.

One of the most common views put forward by supporters of pure vision (especially Tesla) is:

Humans do not have LiDAR or millimeter-wave radar, and they can drive only relying on their eyes and brain.

Dolgov recognized part of this logic. He said that of course a large number of driving tasks can be completed only by cameras. In fact, most of today's assisted driving systems rely on cameras to complete environment perception most of the time.

However, the problem is that assisted driving and driverless driving face two different safety goals.

In L2 assisted driving, the driver is still present. When the system encounters a situation that cannot be handled, it can remind the driver to take over.

But in L4, such as in Robotaxi, no one in the car is responsible for the final judgment, and the system needs to handle all situations by itself, which means it needs to reach a safety level far higher than that of ordinary human drivers.

Dolgov believes that the core problem faced by the pure vision route is that with the continuous improvement of capabilities, the performance curve may gradually slow down, reach its peak prematurely, and it may be difficult to achieve real driverless driving.

In the early stage, by adding data and training larger models, the system capability improves significantly. But when the goal shifts from "approaching human driving" to "far exceeding human driving", the information obtained only by cameras may become a limiting factor.

In short, cameras can help machines see the world, but the information obtained by the cameras themselves has boundaries. The model can improve the understanding ability, but it cannot completely make up for the lack of input information.

Corresponding to Tesla's business, this is also an indirect judgment that Tesla can achieve FSD, but its capability may be limited when doing Robotaxi.

The Route Dispute Between Waymo and Tesla

Dolgov also showed some Waymo test cases in his speech, including night roads, low-light environments and complex traffic scenarios.

However, the key point these cases want to convey is not that LiDAR can replace cameras. In fact, Waymo also relies on cameras to obtain visual information such as colors, texts and traffic signs.

The core he wants to convey is that in the multi-sensor fusion route, different sensors degrade in different ways when facing environmental changes:

  • Camera

Relies on light;

  • LiDAR

Emits laser actively, which can directly obtain the distance and 3D structure of objects;

  • Millimeter-wave radar

Can provide speed information and maintain certain capabilities in environments such as rain, snow and fog.

Therefore, through the combination of multiple sensors, the system can reduce the risk caused by the failure of a single perception method.

For example, in environments such as night, backlight, bad weather or complex occlusion, the vision system needs to infer the distance, speed and spatial structure of objects based on limited information.

The information provided by LiDAR and millimeter-wave radar is of another dimension: one is responsible for directly measuring spatial distance, and the other can provide speed information. Their value lies in reducing the part that the system needs to "guess".

In Dolgov's view, in safety-critical fields, the information that can be directly measured should not be completely handed over to the model for inference.

This is also the biggest difference between Waymo and the pure vision route. The long-standing route debate between the two companies is superficially focused on LiDAR, but the deeper difference lies in the different path design for the development of autonomous driving of the two companies.

Since the initial stage of the project, Waymo has designed the system around driverless driving. The vehicles can operate in limited areas, and can cover a few cities first, but after entering the operation scope, the driving task must be fully undertaken by the system.

Therefore, the priority problem that Waymo needs to solve is how to make a car without a driver operate safely within a limited range.

This determines that from sensors, computing platforms to test systems and operation modes, everything is built around L4.

Tesla adopts another route: FSD is first installed on mass-produced vehicles.

The system expands the data scale through a large number of user vehicles, and then continuously improves its capabilities through model training and software updates, and finally crosses to the driverless driving stage.

The biggest advantage of this method is data scale and iteration speed. Millions of vehicles are running on real roads at the same time, which can provide a data volume that is difficult for traditional Robotaxi test fleets to achieve.

But at the same time, the current FSD is still a system that requires driver supervision, and the driver still bears the ultimate responsibility.

Therefore, the real issue debated by the two companies is focused on whether a set of capabilities accumulated mainly relying on driver supervision can continue to improve along the same technical route and finally reach the safety level required for driverless driving.

However, it is too early to draw a final conclusion on the success or failure of the routes, and we still need to wait for the verification of actual operation data over a longer period of time.

What else did the Waymo CEO say in the 50-minute speech?

In addition to the pure vision controversy, Dolgov also shared some key thoughts of Waymo in the process of autonomous driving R&D throughout the speech.

These include the difference between physical AI and digital AI, the foundation models that Waymo is exploring, the application of end-to-end models, and the autonomous driving safety verification system.

These contents also reflect that the current L4 autonomous driving competition is evolving from simply competing for algorithm capabilities to a competition of system engineering capabilities:

1. Physical AI has a lower fault tolerance rate: errors cannot be undone

Dolgov emphasized many times in his speech that there is an obvious difference between autonomous driving and traditional Internet AI.

If a chatbot generates a wrong answer, the user can ask a new question; if the search result is inaccurate, the user can change the keywords.

However, the wrong judgment made by the autonomous driving system will directly affect the movement of the vehicle in the real world, and the system does not have the option of "undo".

In addition, autonomous driving also faces the requirement of real-time performance. During the high-speed movement of the vehicle, perception, prediction and decision-making must be completed in a very short time, and a large number of calculations need to be executed on the vehicle side.

This is also why the development speed of autonomous driving is slower than many digital AI applications. The improvement of model capability is only part of it, and how to make the model run stably in the real world is another completely different set of problems.

Dolgov mentioned in his speech that the "rapid trial and error" often mentioned in the Internet industry is not completely applicable to autonomous driving.

For AI systems in the physical world, safety needs to be considered from the stage of architecture design, rather than being solved after the product scale expands.

2. Waymo is building a multi-modal driving model

In terms of model direction, Dolgov introduced the technical changes of Waymo in recent years. Early autonomous driving systems usually consist of multiple modules:

The perception module is responsible for identifying vehicles, pedestrians and road environments;

The prediction module judges the future behaviors of other traffic participants;

The planning module determines the next action of the vehicle.

With the development of large model technology, more and more capabilities are being integrated into a unified model. Waymo is currently exploring similar directions.

Dolgov introduced that Waymo is developing a multi-modal model that can simultaneously understand the environment, predict behaviors and generate driving actions.

The model input includes camera, LiDAR and millimeter-wave radar data, and the output involves how the vehicle understands the current environment and how to act next.

From the perspective of technical trends, autonomous driving is experiencing changes similar to other AI fields:

In the past, it relied on a large number of manually designed rules and module combinations, but now it increasingly relies on the data-driven capability of large models.

3. The driving model has both fast path and slow path

Dolgov also introduced a design idea in Waymo's driving model: fast path and slow path.

The fast path is responsible for handling real-time, high-priority driving events.

For example, a pedestrian suddenly crossing the road, a vehicle in front suddenly changing lanes, an obstacle suddenly appearing, etc. These situations require the vehicle to respond quickly to reduce decision-making delay.

The slow path is responsible for more complex understanding. For example, the vehicle sees a smoking car parked on the side of the road.

From a geometric point of view, the road may still be passable. But the system needs to understand that "vehicle on fire" means there may be subsequent risks such as people, fire vehicles or road changes.

Therefore, the driving model not only needs to identify what is happening in front of it, but also needs to understand the meaning behind the event.

This kind of autonomous driving system needs to have two capabilities at the same time, which is actually similar to the "rapid response" and "complex judgment" in human driving.

4. How does Waymo view end-to-end?

End-to-end autonomous driving has become a hot spot in the industry in the past two years. Companies such as Tesla, Wayve and Xpeng are all promoting large models to directly generate driving actions from perception inputs.

Dolgov also talked about end-to-end in his speech, saying that Waymo recognizes the capability improvement brought by large models, but has not completely adopted the "black box" end-to-end route.