GPT Images 2.5 has been unexpectedly launched, and the "soul paintings" created by netizens are each more absurd than the last.
Just now, OpenAI released its new-generation image generation model ChatGPT Images 2.5.
The official announcement states that the new model has been optimized with key improvements in image detail quality, generation speed, editing precision, and consistency across multi-round modification processes. It also adds new features including Sketch hand-drawn reference, creation templates, and image comment editing, enabling users to complete AI creation in a workflow much closer to professional design procedures.
OpenAI notes that currently, users generate more than 3 billion images every week through ChatGPT Images and the GPT Image series models in its API.
The rollout of Images 2.5 reflects OpenAI's intention to further expand the usage scenarios of AI image tools in personal creation, marketing design, product visualization and other related fields.
Compared with the previous generation Images 2.0, the biggest change of Images 2.5 lies in its significantly improved understanding of "modification".
In the past, AI image generation tools were generally better at generating images from scratch, but when users asked to adjust a certain detail, the model would easily modify other unrelated parts by mistake. For example, when modifying a character's clothing, it might alter the facial features; when adjusting the background, it might cause the overall style to change unexpectedly.
OpenAI says that Images 2.5 has been optimized for precise editing, which can execute user-specified modifications more accurately while retaining the original subject, composition and visual style.
Users can adjust individual elements such as products, backgrounds, and text without regenerating the entire image. It truly edits exactly where you point it.
In the demo cases presented by OpenAI, a portrait photo can be converted into different scenes and styles, for example, turning an ordinary character avatar into a retro studio portrait, or converting a pet photo into a new visual theme, while retaining the key features of the original character and animal.
I have also actually tried multiple different types of generation tasks with Images 2.5.
For example, in the portrait generation test, I asked the model to retain the facial features of the reference character and generate a high-definition creative photography work.
The actual test results show that Images 2.5 has an improved ability to retain the features of the reference character. The model can maintain a high consistency of the character's identity while changing the photography style, clothing and environment.
Moreover, Images 2.5 can now finally display the image generation progress intuitively with numerical indicators.
According to tests from netizens, users can even play the Snake game while the image is being generated.
Similar capabilities are also reflected in the scenario of film and television character synthesis. For example, I tested the scenario of Elon Musk "taking a selfie with The Godfather on the film set", asking the model to retain the facial structure, expression and posture of the characters in the reference photo, and generate a 1:1 ratio behind-the-scenes movie selfie.
This type of task used to easily cause problems such as changes in the character's face and wrong proportions, but Images 2.5 maintains the character subject more stably.
Under the street photography style, the "life-defining photo" of Elon Musk is also produced.
In addition to character-related generation, Images 2.5 also has enhanced understanding of product visual design. In a product design test, I asked the model to generate a cola can made entirely of soft plush material.
In the complex visual logic test, I tried to generate a recursive image.
The prompt requires "An orange cat sitting on an office chair holding an iPad, the screen of the iPad shows the same cat holding the same iPad, and the loop continues endlessly".
This recursive structure requires the model to understand the subject, the screen content and the repeating relationship at the same time. The generated result can recognize this visual nesting relationship and achieve a multi-layer recursive effect.
In addition to realistic images, Images 2.5 also has an improved understanding of complex art styles.
In this image of Black Myth style, the overall style of Images 2.5 is closer to the original concept art of the game, rather than simply piling up related symbols.
There is also an illustration in the comic strip style of *Romance of the Three Kingdoms*.
The generated results can understand the characteristics of traditional comic strips such as black and white line drawing, light color painting, and old paper texture. The overall visual effect is closer to the classic Chinese illustration style, rather than simply adding ancient costume elements.
Just like the previous generation, Images 2.5 performs outstandingly in Chinese text rendering.
Overall, the improvements of Images 2.5 are mainly reflected in three aspects.
First, its ability to understand complex prompts is enhanced, and it can process multiple styles, subjects and constraints at the same time. Second, the reference image retention capability is improved, and core elements such as characters and products are more stable during multiple modification processes. Third, the editing process is closer to real design work, and users can continuously adjust around the same image instead of generating new results repeatedly.
According to the actual test results from netizen @aimlapi, compared with Nano Banana Pro, Images 2.5 delivers results with significantly more distinct artistic styles.
However, Images 2.5 is not without shortcomings, for example, the noise performance can still be described as disastrous.
In addition, some netizens found that there are still some detail flaws in the official demo cases of OpenAI.
In addition to the improvement of model capabilities, OpenAI has also added a series of new features around the creation workflow for ChatGPT.
Among them, the Sketch feature allows users to draw sketches directly in ChatGPT and use them as visual references for image generation. Users can simply draw the room layout, clothing outline or creative draft, and then let the AI generate a complete image according to the description.
For example, I can generate a notebook and a pen with a casual doodle.
Sketch intuitively expresses the picture layout and creative intention through hand-drawing, solving the problem that users find it difficult to accurately describe spatial relationships with text and adjust prompts repeatedly.