HomeArticle

The people who know the iPhone Portrait Mode best want to give GPT a pair of "photographer's eyes".

爱范儿2026-09-21 08:12
In the future, cameras will become the basis for machines to judge their next actions.

Over the past week, two new mobile phone products have been going viral online. One is the newly launched iPhone 18 from Apple, and the other is the consumer version of Doubao Mobile Assistant, whose first preloaded device, the Nubia NaviX Ultra, goes on sale today. AI is stepping into mobile operating systems, and we have also published a detailed review of it.

AI has begun to see the screen, and also started to reach out and touch it. Meanwhile, *The Wall Street Journal* revealed that OpenAI has acquired camera technology company Glass Imaging for more than 300 million US dollars.

The transaction has not been publicly confirmed by OpenAI yet, but the position where this new puzzle piece appears is very subtle: Last year, after OpenAI spent 6.5 billion US dollars to acquire io, the hardware company founded by Jony Ive, no major achievements have been made so far, but industrial designers and camera engineers have already joined the team one after another.

It's AI, but not photo-editing AI

Glass Imaging was founded by Ziv Attar and Tom Bishop, both of whom previously worked on computational photography at Apple and participated in the development of the iPhone Portrait Mode.

However, what Glass Imaging researches is not the most common AI photo editing we see today, but using AI to reversely influence the design logic of camera hardware.

After you press the shutter on a mobile phone, what the sensor obtains first is not a directly viewable photo, but a bunch of raw light signals. The traditional image signal processor needs to complete processes such as noise reduction, color calibration, and sharpening before generating the final image on the screen.

Glass Imaging tries to take over this entire pipeline with a set of neural networks trained for specific lenses and sensors. Its underlying technology is still Neural ISP. Traditional ISPs will sequentially hand over RAW data to modules such as demosaicing, noise reduction, multi-frame fusion, and sharpening, each step processed by an independent algorithm, and part of the information will be lost in the process.

GlassAI starts from the RAW burst data of the sensor, puts demosaicing, noise reduction, deblurring, and multi-frame fusion into the same jointly trained network, and at the same time adds lens aberration deconvolution and sensor crosstalk correction that traditional ISPs are not good at, finally directly outputting the finished RGB image.

Based on Neural ISP, Glass Imaging has absorbed and rewritten a large number of core links of traditional ISPs. What it needs to learn is how a camera generates errors, and then reverse-correct those errors when the image is formed. The idea of Glass Imaging is that as long as the lens delivers sufficient information to the sensor, the remaining blur, chromatic aberration and distortion can be repaired by a specially trained network.

But the so-called "repair" here does not refer to the "editing" in the sense of AI photo editing. Generative photo editing can indeed "guess" the blurred text, the moon and human faces in the distance to be clearer, but clarity does not equal authenticity. According to Glass Imaging's materials, its network is trained for specific optical systems and RAW data, with the goal of restoring the information actually captured by the sensor, rather than generating new textures based on common sense and inference.

This is very different from the current main role of AI in image processing, the reason being: if the camera is only used to take photos for social media Moments, it is enough for the photos to look good; but if the image is to be handed over to an Agent to judge the reality and take actions, the objects and text generated out of thin air may lead to wrong operations — this is why we can understand why OpenAI is interested in it.

Two Types of "Vision Rights" for AI Devices

For mobile phone manufacturers, Neural ISP technology allows smaller lenses to capture clearer photos; but for a company that is manufacturing AI hardware, it has another meaning: the world the model sees first depends on what the camera delivers to it.

AI devices need two types of "vision" capabilities. The first one happens inside the screen: AI needs to understand what the user is browsing, which local information is being read, and obtain the permission to call different Apps. On this basis, it can understand the context of the entire operating system.

The second type happens outside the screen: the text on the menu, the objects on the table, the roads in front of you, and the devices in the room cannot be "known" by AI only through chat records. You need to take a photo first, and then submit the photo to AI.

A photo is the end point of human viewing, but it may only be the starting point of an Agent's work. If future AI devices want to take actions in a timely manner, they must shorten the entire chain of "human seeing, shooting, uploading, and querying".

The identity of the camera has therefore changed. In the past, it helped people save frames worth looking back at. In the future, it will become the basis for machines to judge the next action to take.

GlassAI's zoom technology has been applied to the Honor 600 series, which means it has at least completed the adaptation from models, mobile phone chips to mass-produced devices; the company has also demonstrated real-time 4K video processing and more than 20x zoom enhancement on the Snapdragon 8 Elite Gen 5 reference device.

These capabilities are currently used to help people capture distant objects more clearly. In the future, when the camera is not connected to the photo album, but to an Agent that is ready to take action, the evaluation criteria for imaging will change.

Photos not only need to look good, but also present text, objects and spatial relationships as real, stable and timely as possible, because the model may subsequently identify products, operate devices, and even plan actions based on these contents.

It's not that GlassAI will automatically turn the camera into an agent, but it fills the most fundamental link before the agent enters the real world: delivering the information actually seen by the sensor to the model as accurately as possible.

The first change revealed by Glass Imaging is that the division of labor between hardware and software has been redefined. In the past, camera manufacturers first used lenses and sensors as much as possible to solve aberrations, blur and noise, and then let the ISP modify the final image; GlassAI allows hardware to retain some defects that can be reversed by calculation. As long as the sensor still captures sufficient information, the specially trained network can restore it during the imaging process. As a result, lenses can be made smaller, pixels can continue to shrink, and part of the problems that originally required adding more lenses, increasing thickness and cost to solve are transferred to the neural network.

So it is possible that the hardware of future AI devices does not need to be fully designed before "inserting a model" into it. Lenses, sensors, NPUs, memory and models will be co-designed from the very beginning. Questions like which raw signals the sensor is responsible for retaining, which errors the neural network can repair, how much local computing power the device needs, and whether it can continuously understand the environment in a low-power state... all these problems will in turn affect the shape of the hardware.

For example, if engineers know in advance that certain optical defects can be stably repaired by the neural network, when choosing the number of lenses, pixel size, sensors and body thickness for the next generation of cameras, they can tolerate defects that were unacceptable in the past. It may no longer be necessary to stack more lenses to completely eliminate aberrations, or it may be to ensure that the sensor still retains enough raw information, and then hand it over to AI for restoration.

The same logic may also extend to microphones, spatial sensors and other perception components. In the past, hardware was responsible for recording the world as completely as possible, and software only processed it afterwards; in the future, hardware only needs to capture the information that the model can understand and restore, and the model will participate in influencing how the hardware should view the world from the very beginning.

The model is not responsible for deciding what the camera sees, but it will change how much hardware the camera needs to obtain an image sufficient for restoration and understanding.

This article is from the WeChat Official Account "ifanr", written by ifanr that discovers tomorrow's products, and published with authorization from 36Kr.