HomeArticle

The most interesting photography-related upgrade of the iPhone 18 Pro is two seemingly "contradictory" AI features.

爱范儿2026-09-18 15:20
In the future, cameras need to prove that "the pixels above come from me" before any algorithm intervenes.

At 8 a.m. today, the Apple iPhone 18 Pro series officially went on sale, with long queues forming at offline stores in many regions. Since the most eye-catching iPhone Duo will not be available until October, the two models launched today, the iPhone 18 Pro and iPhone 18 Pro Max, have drawn relatively moderate public enthusiasm. Especially for the former, the price has dropped below the official guidance level on e-commerce platforms during the pre-order period.

Just three days ago, iOS 27 was finally officially released, and a malfunction immediately came to light. A Reddit user submitted a photo of a horse to Apple's new Spatial Reframing feature, expecting the tool to fill in the missing part of the frame, but a complete stranger unexpectedly appeared in the resulting image.

It was neither the user himself, nor any of his family or friends that had been photographed in his album, but a total stranger, a non-existent "person" that was a hallucination generated by AI.

Left: Original photo, Right: Photo modified by Spatial Reframing. Image from Reddit user @friendofmany

As can be seen, there is a small object that looks like a hat behind the horse in the original image. When Spatial Reframing expands the frame by shifting the view upward according to the image content, it generates a man that has never appeared in the original scene.

This sounds quite risky. How could the iPhone do such a thing? Don't worry, if you plan to buy the iPhone 18 Pro series, it is equipped with the latest Apple Reference Image feature that acts as a countermeasure against AI-generated content.

But hold on a second, does this mean that as iOS with a large number of generative AI features hits the market, Apple has also launched a feature that works in the opposite direction of generative AI?

Yes, this is a set of seemingly contradictory new features we discovered on the iPhone 18 Pro series.

A hat ended up generating a whole person

Spatial Reframing is a new generative image editing feature added in iOS 27. It was first available in the beta version back in June. Users can not only crop photos, but also drag the frame to simulate the viewing angle after the photographer moves sideways.

But since the camera does not actually move, for Spatial Reframing to change the camera position after the shot is taken, it must draw the areas that the original lens did not cover. Where there are no pixels, it generates pixels; where there is a missing person under the hat, it fills in a face and a full body for it.

The principle behind this image technology still relies on generative AI, so "creating something out of nothing" is unavoidable. Testers from TechRadar also found that when processing images of buildings and busy streets, the AI occasionally adds windows and flanges to buildings that do not exist in reality. Compared with ordinary image outpainting, Spatial Reframing does not just fill in more background around the photo. It simulates the movement of the camera position, reprocesses the relative relationship between the foreground and background, and then fills in the areas that were originally blocked or did not enter the lens at all after the viewing angle changes.

Apple can continue to improve the perspective, light and shadow, and object continuity to make the generated results more and more natural, but it cannot eliminate the fundamental contradiction of this feature: as long as the camera does not record those pixels, the model can only guess what should be there based on the existing image.

To be honest, mobile phone photography has never been a process where light passes through the lens and falls into the album completely unchanged. The moment you press the shutter, a series of links work together: HDR merges multiple exposures, night mode brightens the details in the dark, and portrait mode calculates the distance between the subject and the background. The photos we see are always the result of the joint production of the sensor and the algorithm.

However, in the past, these calculations were mainly responsible for making the content already captured by the camera clearer. Spatial Reframing crosses another line: it needs to fill in content for areas that the camera did not see.

As a result, the photo still looks like a normal photo, but it is no longer naturally equivalent to an on-site record. That stranger did not leave any trace of entering the frame, nor was the moment the shutter was pressed captured; he was directly generated from the original small piece of hat into a living "person".

Apple leaves a second piece of evidence for your photos

However, Apple is obviously not unprepared. The brand new iPhone 18 Pro series will be equipped with Apple Reference Image.

This is an optional Reference shooting mode. When enabled, the camera will leave a DNG "digital negative" directly signed by the sensor when taking photos, recording the original pixels, sensor information and shooting time, and bind it to the main photo processed by conventional computational photography.

Image from: Instagram user @iJustine

This verification system is established from the moment the mobile phone leaves the factory. The main camera sensor will generate a unique encrypted identity and bind it to the phone's Secure Enclave. When the user presses the shutter, the sensor enters a special security mode, and directly signs the pixel data before the firmware modifies the image; then Private Cloud Compute checks whether the sensor, mobile phone and timestamp match.

The system will save an editable ordinary photo at the same time, as well as a secure digital negative signed by the sensor. The latter will not immediately become a viewable Reference Image; only when the user needs verification, it will be developed into an untampered reference image via Private Cloud Compute.

When it was first released, it was understood as a kind of "watermark" to identify AI-generated content, but in fact, it provides an independent reference that can be used to compare what happened between the final photo and the sensor record.

Take the photo of the horse's back taken by the Reddit user at the beginning as an example (obviously he is not using the iPhone 18 Pro series). If you use the normal shooting mode and then use Spatial Reframing to change the viewing angle, what will appear in the album is only a complete photo after generative processing. It may retain the editing history of Apple's own album, and you can click "Revert", but there is no sensor signature and no Reference Image for external verification.

On the iPhone 18 Pro series, the Reference mode will not immediately add three extra photos to the album. When you press the shutter, two processing paths start at the same time: the conventional computational photography process generates a main photo that can be further edited; before the firmware modifies the image, the sensor signs the captured pixels and saves it as a secure digital negative associated with the main photo.

This DNG negative is not yet a Reference Image. Only when the user needs verification, it will be sent to Private Cloud Compute to complete demosaicing, tone mapping and compression in an auditable environment, and finally become a viewable reference image with Apple's signature.

If the user later uses Spatial Reframing on the main photo, the album will first display the edited version. The main photo before editing is still retained in the non-destructive editing record, and the Reference Image is retained along another path, which can only be accessed through the Reference mark.

As a result, three levels appear in the same photo project:

The edited final image, with a stranger standing behind the horse;

The photo after undoing the editing, with no one there;

The Reference Image generated from the sensor record (needs to be called up), with no one there.

The first two record how Apple processes the photos, and only the last piece of evidence can prove that the sensor never saw this person the moment the shutter was pressed.

Redundant? Far from it

With today's technology, the concept of "original image" is already a delusion of chasing a mark that has drifted away.

If you do not turn on the Reference mode when taking photos, even if there is an "original image" in the album, all its modifications and edits can still be undone and restored in the Apple album, but there is never a hardware-signed original record that can be handed over to others for verification.

In front of the final image processed by Spatial Reframing (or any generation tool), the "original image" is just material — and these materials can also be completely generated.

Reference Image pushes the starting point of trust to the sensor, so that "original record" is no longer only claimed by the album, camera application or operating system.

After entering the Reference mode, the sensor will restart and enter a dedicated secure shooting state. It converts the received light into pixel data, and completes the encrypted signature directly inside the hardware before the data leaves the sensor. When the operating system receives these pixels, a shooting record that cannot be replaced without a trace has already been formed.

However, the sensor is only the starting point of this trust chain. The subsequent device identity, shooting time and image development need to be verified by Secure Enclave, encrypted timestamp and Private Cloud Compute respectively. If any link cannot correspond to the original sensor signature, the final Reference Image cannot maintain a valid verification status.

Therefore, the Reference mode does not dwell on the concept of "original image". It re-establishes a boundary with the sensor as the core of fact, and builds a set of encryption and verification pipelines around this hard boundary.

Within the boundary is the pixel data formed by the optical signal; outside the boundary, the content may come from PS stitching, supplementary calculation by Spatial Reframing, or another set of AI generation tools.

In the future, cameras will deliver two photos with three states

In the future, photos may no longer only be divided into "original / edited", but there will be three more detailed states:

No Reference Image exists, and the shooting source cannot be proved;