HomeArticle

It turns out that these details are the exact reasons why AI-generated images look jarring and unnatural.

神译局2026-09-24 07:24
The footage lacks causal logic.

Shen Translation Bureau is a translation team under 36Kr, focusing on technology, business, workplace, life and other fields, with an emphasis on introducing new technologies, new perspectives and new trends from overseas.

Editor's Note: Many AI-generated images have exquisite and realistic textures with well-rendered details, but they always make people feel subtly incongruous. The problem is often not the wrong human body structure, but the fact that the picture violates the physical logic of the real world. This article is translated and compiled, hoping to bring you inspiration.

A large number of AI-generated images are highly realistic, with delicate textures, clearly visible skin pores and outstanding overall visual aesthetics. Yet there always seems to be something off. You cannot pinpoint the exact problem, but you are fully aware that it is not a real photo. The root cause is that we are looking for flaws in the wrong places. Instead of focusing on those obvious low-level errors, we have never learned to identify the remaining deep-seated vulnerabilities in current high-level AI-generated images. The hard errors visible to the naked eye in images generated by advanced AI today are very few, and the flaws have become extremely hidden: only results, no causes.

Only Results, No Causes

In real photos, every visual effect on the frame can find a corresponding physical source. This shadow comes from a certain object; this blur comes from a specific aperture and shooting distance; the reflection in the eyeball comes from a certain window. A camera does not create out of thin air, it only records the results produced by real causal relationships.

Generative AI is the exact opposite. It learns what things look like, but does not understand how they work. It knows what reflections look like, what shadows look like, and what portrait blur effects should be, but it does not understand why these phenomena occur. Therefore, AI will randomly draw light and shadow effects at aesthetically pleasing positions, completely ignoring the physical rules that constrain the positions of light and shadow in the real world.

This is the root of all problems. Remember this sentence: AI does not understand causal relationships, it only understands correlations. AI has seen millions of images, and knows that eyes usually have highlights, backgrounds are mostly blurred, and glasses generally have reflections. It only knows which elements often appear at the same time, but does not understand what is the cause and what is the effect. That is why AI can generate pictures with only effects but no corresponding causes. For the model, there is no such thing as a cause at all, only visual elements that often appear in pairs.

This is the distinguishing feature of 2026 AI images: they no longer have obvious flaws that can be pointed out at a glance, but have perfect visual effects yet lack corresponding real-world causes. Once you master this identification method, you will no longer be deceived by AI-generated images. Several real cases are listed below, analyzed one by one from relatively obvious to extremely hidden.

Eyes Are Where Flaws Are First Exposed

When identifying AI-generated images, prioritize looking at the eyes. This is where model errors are most likely to occur, yet almost no one pays attention to it.

Look at the iris in this picture, there is a star-shaped light spot inside the pupil.

The star-shaped highlight in the eye can only come from a light source with a specific shape: a ring light, a partitioned window, or a filter, which is essentially a reflection. The reflection should appear on the cornea, the moist outer surface of the eyeball. The iris is located under the cornea and inside the aqueous humor. The iris itself does not reflect light, it only absorbs and diffracts light.

In this picture, AI wants to draw the reflection of a shaped light source, but mistakenly carves the star-shaped light spot inside the iris, destroying the radial physiological structure of the iris. The highlight is drawn to a lower position that it should not belong to. This star shape is not on the surface of the eyeball, but embedded in the eye tissue. No lighting fixture in the real world can create such an effect. For light to reach this position, it must penetrate a layer of tissue that is supposed to block light.

The iris should be a regular circle. Once the shape of the iris is distorted, it can basically be judged that there is a problem with the picture.

Eyeball Reflections Are Real Snapshots of Light Sources

There is another widely recognized key point in the field of photography: the reflection in the eyeball can restore the real position of the light source.

The shape of the catchlight completely corresponds to the luminous object itself - a square softbox, a rectangular window, a ring fill light. By reading the reflection in the eyeball, you can restore the entire lighting setup. If the light source corresponding to the eyeball reflection contradicts the lighting on the rest of the face, the realism of this picture will collapse directly.

This night selfie taken inside a tent is a typical case.

First, you can see that the iris has been deformed, but this is not the only reason that breaks the photorealistic effect. The lighting logic is also very strange.

The light on the front of the forehead is harsh and strong, which is the effect of a flash, whether it is a top-mounted flash or a handheld flash. A flash means that the camera has a hot shoe interface. Combined with the compressed perspective of the human face, the camera should use a 50mm or longer focal length lens. Here comes the contradiction: the picture is a selfie perspective with arms outstretched, but the device that can achieve this imaging effect is a telephoto DSLR. The length of a human arm cannot reach this shooting distance at all. The device that produces this strong light cannot be held in the hands of the person in the picture.

There is another detail that can only be noticed by people who have been exposed to light metering equipment for years. To use a front flash without making the person look stiff and abrupt against a dark background, photographers use the technique of balancing ambient light and flash. First measure the background parameters, set the camera according to the background exposure, then lower the flash power to make the picture blend softly. AI reproduces the soft gradient picture brought by this balancing technique, but at the same time superimposes an extremely intense flash, and the two cannot coexist. AI copies the visual result presented by the technique, but does not understand the principle behind it.

Midjourney is imitating the finished works of photographers, but does not understand the operational logic of photographers. This lack of cognition is revealed through the lighting effects that cannot exist in reality.

Unattributed Shadow

This flaw is even more hidden, but once you notice it, you can never ignore it again.

There is a shadow on the shoulder and chest of the T-shirt. This shadow indicates that there should be a wall or object on the side blocking the sunlight. But if such an obstacle really exists, the face, which is at a higher position and on the same plane, should also be covered by the shadow. However, the face is in a fully lit state.

The shadow stops abruptly at the hairline with a sharp edge, but there is no solid entity that casts the shadow. Even if you forcefully imagine a raised structure to explain the shadow, it does not make sense from the perspective of photographic composition: this shadow neither divides the frame nor sets off the atmosphere, and has no compositional function at all. Its appearance violates both physical rules and creative logic.

The underlying reason is that the lighting of the character subject and the lighting of the environment are generated separately by AI, and the two sets of lighting do not match each other. In a real photo, the sun is a unified light source, and the shadows on the face and objects are all constrained by the same set of light, no additional adjustment by the photographer is required. AI does not have the concept of a unified "light". It separately learns the two image samples of "bright face" and "shadow on cloth", and forcibly stitches the effects together, ignoring the physical rules that require the two to be self-consistent. This makes the character and the background look like two photos pasted together.

In a real scene, the same set of light sources illuminates the entire frame. If a picture has two sets of conflicting lighting logics, it is essentially a collage of two pictures.

Flat Blurred Background

This is the flaw with the highest recognition rate, yet the most easily ignored by the public. Because background blur is equated with professional photography in the public's perception.

Mobile phone selfies have two characteristics: distortion caused by wide-angle lenses, and a large depth of field where almost all objects are sharp. This picture meets neither of these conditions. It has the blurred background of a portrait shot with a large-sensor camera, but the composition is a selfie perspective with arms outstretched. There is no camera equipment in reality that can achieve both effects at the same time.

If you observe the blur itself carefully, the flaw becomes even more obvious. There is a rule for optical blur: the degree of blur increases with distance. An object one meter away is less blurred than an object five meters away. But in this picture, the bricks, lamps and curtains are all blurred to exactly the same degree, as if all the background objects are at the same distance. This is a fake 2D flat-style blur, just like blurring the entire image in Photoshop and pasting it behind the character.

This is no coincidence. The AI training dataset contains a large number of post-production composite materials: cutout character layers superimposed with blurred 2D backgrounds. AI learns this opportunistic pattern and reproduces it continuously. If the background blur of the picture has no gradient according to distance, it means that AI is calling these low-quality composite samples from the training set.

Sample picture: Under a large aperture, the degree of blur gradually increases with distance. Image source: Sasun Bughdaryan, Unsplash

Look at the glasses worn by the character, there is another hidden flaw - no nose pads. The small component on the frame that is supposed to rest on both sides of the bridge of the nose is completely missing. The lenses and frames look realistic, but AI does not understand the mechanical structure of how glasses are worn, so it omits small functional parts such as nose pads, hinges, and temples that enable the glasses to stay in place.

The glasses have no nose pads.

Real out-of-focus blur exists in the depth of field; fake blur is just a flat backdrop. Pay attention to another detail: the lenses have no reflections at all, just like empty frame props. The next case will also have this problem.

Sample picture: AI-generated image after optimizing reflections and depth of field effects

Glasses Without Lenses Are Just Props

This is not an accidental error in a single image. The model will repeatedly make similar mistakes, which can be confirmed by the second case.

The lens is partially missing, and the iris is also slightly deformed.