HomeArticle

Why is the 1-megapixel camera the most important component for Apple's AI?

爱范儿2026-08-24 09:01
Build a user-owned personal context with the Apple full product suite.

A few days ago, MacRumors unearthed a pre-made Apple promotional video directly from macOS Tahoe 26.7 RC: you wear AirPods, pick up a book, Visual Intelligence identifies the book title, and then Siri is asked to save it.

The product code name in the video is B790. According to Mark Gurman, it was already on Apple's 2026 product roadmap, and the subsequent mass production version B798 will not be released until 2027.

The recently unearthed system code also lets us see for the first time exactly how these headphones "see". The most interesting part is a seemingly underwhelming parameter:

The camera on AirPods only has 1 megapixel.

Photo | WIRED

At a time when the mobile phone industry is already discussing 200-megapixel sensors and 8K video, the camera Apple prepared for its future AI hardware may actually capture images at a resolution as low as 320×320.

It is obviously not designed for taking photos.

Why does the camera on AirPods only have 1 megapixel?

Judging from the currently unearthed code, this camera system has at least two working modes.

In active mode, the camera can capture 640×640 images, with a maximum output of 1024×1024 after processing. In passive mode, it captures 320×320 images, and outputs 320×320 or 512×512 images.

The code shows that the left and right AirPods can capture RGB still images synchronously, align the two perspectives through the same Frame ID, and then continuously collect data at a certain frequency for Visual Intelligence to process.

The system processes ambient audio, changes in surrounding sounds, posture and head rotation information at the same time; the two cameras need to be calibrated, and the camera and motion sensors also need to be calibrated. If the head movement is too violent or the lens is blocked, the corresponding frame will be discarded directly.

Code related to "peripheral inference" even appears in AccessorySensorManager, which can judge whether there is a person in the frame on the device side.

All of these are perfect data fed to AI.

For AI, in most cases, 48-megapixel photos are not needed at all. The model only needs to know that there is a book on the table, a restaurant ahead, a drink in hand, and what is written on the road sign.

A resolution of 320×320 is already sufficient.

As the resolution decreases, the amount of data is reduced, and the burden of image processing and transmission is also lowered. For a device that has a battery as small as a headphone and needs to be worn for a long time, power consumption and battery life are more important considerations than image clarity.

Photo | iFIXIT

Over the past two decades, cameras in consumer electronics have mainly served human users. Pixels, sensor size, dynamic range, and color performance all point to the same goal: taking high-quality, good-looking photos.

But AI cameras first serve machines.

Photos don't even need to be saved. Once the machine understands the content, the data can be deleted.

Apple is not the first company to develop products along this path. The Light Sail AI full-sensory headset launched this year is a direct reference. It also puts cameras in the left and right ear cups, uses two 2-megapixel fixed-focus cameras, and its core use is not photography, but scene AI recognition.

The real world is first converted into visual data by the camera, and then handed over to the model for understanding. Light Sail even puts 4G eSIM and GPS into the charging case, so that the device can reduce its dependence on mobile phones.

Whether it is 2 megapixels or 1 megapixel is not that important here. The engineering problem has become: can it see clearly, how long does the processing take, how much data needs to be transmitted, and how much power will be consumed.

The earlier Humane AI Pin also tried a similar idea, using a camera to add visual capabilities to AI.

But Humane needs to convince people to wear an extra new device on their chest just for AI.

Apple does not have this problem.

Many people already wear AirPods every day. It does not need to reinvent AI hardware, but only needs to give the existing hardware new sensory capabilities.

In the AI era, AirPods are not just headphones, but sensors

In fact, the two AirPods have already formed a complete sensor suite:

The camera observes what is in front of you, the gyroscope records the direction your head is facing, the microphone collects surrounding sounds, and the position and motion sensors supplement information about where you are and what you are doing.

For example, you walk into a bookstore, pick up a book, and then ask Siri: "Is this worth buying?"

If you only send this sentence to today's large language model, it actually knows nothing at all.

Which book does "this" refer to?

But if AirPods just saw the book cover, knows your head is facing the book, and your iPhone knows you are in a bookstore, then Siri can truly understand what "this" exactly means.

If the system also knows that you have read two books by the same author in the past, finished one and only read a third of the other, it can even give an answer that is meaningful only to you.

The changes that take place here are very subtle.

When we communicate with AI today, we are constantly explaining ourselves to it. Who I am, where I am, what I just saw, what I am holding, and why I am asking this question all of a sudden.

Essentially, most prompts are manually supplementing the missing context for the machine.

What cameras, microphones, position and motion sensors do is to gradually eliminate the need for these manual explanations.

Large language models already know a lot about the world, but they still know very little about "you".

Personal AI should not act like it is meeting you for the first time every time it is woken up. What camera-equipped AirPods complement is exactly a kind of "personal context".

Photo | Apple Insider

Taking photos with a mobile phone is a very deliberate action: pick up the phone, open the camera, and point it at the target.

But headphones are different. They are always worn on your head, so their interaction with the surrounding environment is particularly important.

There is also a small detail in the system code unearthed this time:

Each AirPod seems to be able to control the hardware indicator light, including switching it on and off and adjusting its brightness.

If this corresponds to the working state of the camera, then when AirPods captures environmental images, people around may be reminded by the light — which is very consistent with Apple's consistent privacy protection strategy.

If the camera only takes a photo once in a while, a flashing light may be enough; but if it needs to continuously understand the environment at a low frequency, the problem becomes more complicated.

Because what AI needs is not a single photo, but context, it needs to keep all its senses turned on, and proper reminders are very important at this time.

From this perspective, the positioning of AirPods has also changed subtly — it is no longer just a pair of headphones, but a new form of AI hardware.

What if changes like this do not only happen to AirPods, but appear on more Apple products?

I think this is Apple's most important strategy in the AI era. And the most important time node in this strategy is 2027.

2027 may very well be the first year of Apple AI

2027 is already a special year for Apple — it marks the 20th anniversary of the iPhone's release.

According to current leaks, Apple is preparing the 20th anniversary iPhone with a major design overhaul, and camera-equipped AirPods may also debut around the same time.

At the same time, Apple has several other AI hardware product lines in development.

In addition to camera-equipped AirPods, there are smart glasses, and AI pendants that can be clipped to clothes or worn as necklaces.

In previously exposed product plans, an Apple Watch equipped with a camera also appeared, which gives the watch the ability to observe the surrounding environment.

Viewed separately, these products do not have much connection with each other.

Headphones are headphones, glasses are glasses, watches are watches.

But if you temporarily put aside the concept of "products" and only look at the sensors on them, you will see a different picture.

Many nodes of the context network are already on our bodies today.

The Apple Watch on your wrist knows what time you fell asleep last night and how many steps you walked today; the iPhone in your pocket knows what time you left the office and which restaurant you went to; the Mac knows which documents you opened in the afternoon; AirPods are almost always attached to your ears.

What Apple is preparing to do now is to continue adding sensing capabilities to these devices.

In the past, this information was scattered across different devices and apps. Heart rate belongs to "Health", photos belong to "Photos", location belongs to "Maps"...

After the emergence of AI, for the first time, all of them can become the same thing: context.

Thus, AirPods knows what it hears, and may also know what you are facing; Apple Watch knows what is happening to your body; iPhone knows where you are and who you are communicating with; Mac knows what work you are processing. If smart glasses are finally launched, it will also get a direct first-person perspective.

None of these devices alone is enough to truly understand a person, but each of them holds a different facet of the user. Once they work together, they can gain a panoramic understanding of a person.

This is also the irreplicable difference between Apple and many AI hardware startups.