HomeArticle

Just now, GPT-Image-2.5 was released.

新智元2026-09-09 08:11
Chinese garbled characters no longer exist.

Just now, OpenAI released GPT-Images-2.5:

Generation latency is reduced by up to half, editing accuracy is improved, a new Sketch hand-drawn input feature is added, and for the first time, two tiers of models with different speed levels are separated on the API side.

https://x.com/OpenAI/status/2097394956457623964

It is open to all users of ChatGPT, ChatGPT Work and Codex, including the free version.

Calculated based on this figure, about 430 million images pass through this system every day on average, and the sensory change brought by the accelerated generation covers a far wider range than ordinary feature iterations.

OpenAI announced the total generation volume of ChatGPT Images and GPT-Image API: 3 billion images per week.

Amazing Demo: No more garbled Chinese characters

How powerful is GPT-Image-2.5 exactly? Let's take a look at the Demo directly.

The garbled Chinese characters have completely disappeared!

This is the reference image:

This is the output retro-style image:

Can you still tell the difference between real photos and AI-generated images?

Extracting and restoring old printed photos into high-definition electronic photos is a piece of cake.

Want to try on different outfits to see the visual effect? Just leave it to ChatGPT:

Input:

Output:

Input original image with messy quilt:

Output image with the quilt neatly folded:

It features extremely strong scene consistency, with no visible flaws or mismatched details.

Twice as fast, no unintended changes after editing

The speed improvement is the most straightforward. According to OpenAI, the generation latency is up to 50% lower than that of Images 2.0.

The rendering quality of photos and textures is also improved at the same time, and the restoration degree of the main character in the reference photo is also higher.

Precise editing targets a long-standing problem of image generation tools: when users ask to modify a local part of the image, the model will also change the parts that are not supposed to be modified.

According to OpenAI, Images 2.5 can only modify the specified area while retaining the composition, lighting and main features of the rest elements of the image, with significantly improved stability in complex backgrounds and multi-subject scenes.

Users can also directly place comment annotations on the image to circle the position that needs to be modified, and the operation logic is similar to the design tool Figma.

The consistency of multi-round editing is the third improvement.

When repeatedly modifying the same image in a long conversation, the editing results of previous rounds can be retained, and the image quality will not degrade with the increase of editing rounds.

Axultan Alimkulov, Head of Product at Higgsfield AI, commented on Flare:

What impresses us most is its understanding of "what should not be modified".

Making a meaningful edit will not lose the character, composition and visual style of the original image.

Sketch and templates: from typing to drawing

In terms of product features, Images 2.5 brings four new functions.

The entry of Sketch is to enter @Sketch in the ChatGPT dialog box to open the hand-drawn panel. Users draw a sketch with text descriptions of style and details, and ChatGPT will generate the finished image with the sketch as the composition reference.

The examples given by OpenAI include turning room layout sketches into interior design renderings, and turning character outline sketches into illustrations. The threshold does not lie in drawing skills, but in drawing the spatial relationship and approximate proportion to let the model understand the structure of the picture.

The templates cover high-frequency formats such as posters, product images, and flyers.

Select a template, fill in the information and style preferences, and the model will generate images according to the preset structure, eliminating the trial and error of writing Prompts from scratch.

Prompt sharing allows users to send out the complete Prompt of the generated image, so that others can bring in their own photos and details to generate a personal version under the same framework.

A Prompt themed "1980s retro portrait" has been circulating on social platforms, and users can upload their selfies to get photos in the same style.

The four functions point to the same change: the process of turning the picture in your mind into an image no longer relies solely on typing!

Flare and Sunburst: API is split into two for the first time

For developers, OpenAI simultaneously launched two image models on the API side: GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst.

Flare focuses on speed and batch generation, and OpenAI positions it as the default choice for most applications.

Applicable scenarios include social content, product experience images, rapid prototyping and high-frequency image generation.

Data from Lucky Liao of the Manus evaluation team is more specific: in their tests, Flare's image generation speed reaches 2 to 4 times that of GPT-Image-2, and the improvement of transparent background generation is directly useful for programmatically generating brand materials, presentation slides and web page illustrations.

Sunburst is positioned for high-end creative workflows, which uses longer generation time in exchange for more precise editing control.

The typical scenarios given by OpenAI are finished-grade materials for brand marketing and retouched product images.

Adobe has confirmed that it will integrate the Images 2.5 model into Firefly.

Matt Chotin, Senior Director of Product at Adobe, said that the 2.5 version brings faster generation speed and resolution consistency, and the image can remain clear and realistic after multiple rounds of retouching.

The split of the two models corresponds to two types of demands: high frequency with low latency and high-precision customization.

Flare is for scenarios that require large output volume, and Sunburst is for scenarios that need high-quality finished products. Developers can choose models according to their business scenarios, without paying extra waiting time for unused high precision.

This is the first time that OpenAI's image API has carried out product stratification according to speed and quality, and the logic is consistent with the tiered design of the language model API.

The naming of Images 2.5 continues the rhythm of OpenAI's image product line: the architecture remains unchanged, and the speed and accuracy are greatly improved in this iteration.

GPT-Image-1.5 at the end of 2024 adopted the same strategy.

The two .5 versions are less than a year apart, and the iteration frequency of image models is catching up with that of language models.

This is the most powerful image generation model in the world, with an overwhelming leading edge.

OpenAI's continuously enhanced multimodal capabilities are very likely to help it enhance the capabilities of base models such as GPT-7 by strengthening capabilities including Computer-Use capabilities when competing with Anthropic's models.

References:

https://openai.com/index/introducing-chatgpt-images-2-5/

This article is from the WeChat Official Account "AI Era" (ID: AI_era), author: ASI Apocalypse, editor: Ma Ke, published with authorization from 36Kr.