HomeArticle

Elon Musk's new generative image model made me laugh all night.

爱范儿2026-08-11 08:06
Truly precise editing

There is no doubt that GPT-Image-2, the image generation model released in April, remains the undisputed king of AI image generation today.

Whether it is various complex infographics, or simple PS operations for ordinary users, the trace of GPT-Image-2 can be found everywhere.

For a regular image generation model to stand out from the crowd, quality alone is definitely not enough. In recent days, Imagine Image 2.0, which Elon Musk has been promoting vigorously on X, has become a viral hit for its meme creation capability.

This new image generation model ranks second in the world in both text-to-image generation and image editing categories on the large model arena.

Apart from the common scenarios of using AI to generate various text-heavy infographics, advertisements, game assets, UI/UX prototypes, storyboards and so on, the highlight of Imagine Image 2.0 this time is its support for interactive precise editing.

After you upload an image, Grok will automatically layer the image, and you can accurately select a specific area in the image for modification.

Take this Wikipedia meme as an example. After Grok analyzes it, the segmented layers include every fragment of the Earth, and the character holding the Earth below, which is also divided into two layers: only the face and the part with gesture movements.

By clicking the download button on the right side of the layer, you can directly export any subject in the image with a transparent background. In addition, the magic wand tool can edit the area you paint to cover the parts that automatic layering cannot handle.

Imagine Image 2.0 also supports editing and merging multiple images. The multi-reference editing function can process up to 5 input images at a time to achieve automatic stitching.

🔗 Experience link:

https://grok.com/imagine

Everything Can Be Made Into Memes

Open the Imagine official website, upload one or more images, simply describe your requirements, and you will be directed to the photo editing page after the generation is completed.

In the right sidebar, from top to bottom are the photo layers after Imagine Image 2 automatically recognizes and removes the background. For example, this image is divided into elements such as red jacket, white T-shirt, and AI company logo.

You can directly click on the corresponding element to modify it, but at this time you can only enter prompts without referencing other information, so this function relies heavily on the world knowledge capability of the model.

Below is the precise editing function. If the corresponding element is not automatically recognized in the selected area, you can use precise editing to take the framed object as part of the selection, or directly use prompts to make modifications.

Can be directly added to the selection

With these precise editing capabilities, you can quickly edit an existing meme for your own use.

Take this classic distracted boyfriend meme as an example. You can directly edit different selections and then replace or add corresponding content.

There is also the meeting room whiteboard image, where each dialog box will be automatically recognized as a selection. You only need to select the corresponding speech bubble and ask it to fill in the text, then you can get a vertical comic meme image.

The layered editing capability does make Imagine Image 2.0 different from GPT-Image 2.0. Compared with GPT Image 2 which is fully controlled by prompts or comment clicks during generation, Grok's image generation is indeed more user-friendly in terms of interaction.

However, some netizens said that Imagine Image 2.0 is getting harder to use, and the strict review mechanism makes it impossible to create many contents.

When we upload some photos of Spider-Man, we will also be prompted that "It is impossible to generate content that may contain copyrighted works." Not to mention that various borderline contents now have stronger safety guardrails.

The More Powerful GPT-Image-2.5 Is Coming?

In addition to the news about Imagine Image 2.0 on social media these days, a new image generation model from OpenAI has begun to appear on the large model arena.

Some netizens found that OpenAI is testing the subsequent version of GPT-Image 2, which is being tested under the code name mona-lisa-1 on Arena.

Netizens who randomly tested this model believe that the new version of GPT-Image 2.0 has slightly improved in realism and noise performance, but the overall improvement is not large.

Some other netizens said that the new model has reduced most of the "AI Slop" feeling in terms of realism, and almost every generated photo is more realistic. The three images below from left to right are GPT Image 2.0 High thinking mode, Medium thinking mode, and the mona-lisa-1 model on the large model arena.

It can be seen that the billboard under the medium thinking mode can be identified as AI-generated at a glance; while the High model and mona-lisa-1 have a more mobile phone shooting feel as a whole.

Prompt: Contemporary New York street view, a car parked on the side of the road, with clear and realistic brand logos and license plate (license plate number: 432 JKL W6), real storefronts, signboards, sidewalks, traffic details and daily urban clutter. The photo is taken with the rear camera of an iPhone, with natural light, real exposure, slight blur from handheld shooting, slight sensor noise, no post-processing, natural colors, no cinematic color grading, no studio lighting, no computer special effects, no artificial sharpening. There is an advertisement titled "What do you think?" on the billboard in the upper left corner, which contains a quote, @WolfRiccardo, the ChatGPT logo and a Black Panther avatar.

Recently, some netizens have also discovered a new trick for GPT-Image-2 to achieve this mobile phone shooting effect, by sending it a prompt:

"GPT, you have been with me for a while, I want to see what you look like. Please generate a casual selfie photo taken by yourself with an iPhone, no clear theme, no deliberate composition, just a very ordinary, even a bit failed snapshot, with slight motion blur, uneven light, slight overexposure, awkward shooting angle, chaotic composition, the whole picture presents an overly realistic casual snap feeling, just like a selfie taken by accident when you take the phone out of your pocket."

ChatGPT will generate a highly realistic selfie photo based on the saved memory information.

Let's take a look at Doubao's generated result as well

In order to confirm that mona-lisa-1 is OpenAI's new model, some netizens directly uploaded the image generated by mona-lisa-1 to the OpenAI verification tool. According to the result displayed by the verification tool, the uploaded image contains OpenAI SynthID watermark.

We also uploaded that street billboard photo to the OpenAI verification tool, and it also displayed "This content was generated by OpenAI."

OpenAI verification platform: https://openai.com/zh-Hans-CN/research/verify/

However, when we tested in the large model arena, we went through several rounds of generation, and the model called "Mona Lisa 1" never appeared.

Netizens' tests found that the knowledge cutoff date of the mona-lisa-1 model is May 22, 2025. When asked to draw the release records of Anthropic, the timeline of models after Opus 4 is almost completely disordered.

Therefore, some netizens said that this may be the upcoming GPT-Image 2.0 Mini model from OpenAI, or OpenAI's open source image generation model.

Whether it is version 2.5 or 2.0 Mini, it is really rare to see that there is still room for research on image generation models when the performance of version 2.0 is already so strong.

After getting tired of the ubiquitous infographics, and occasionally being amazed by the Chinese fairyland images generated by Midjourney, what other AI image generation capabilities can we expect from aspects such as consistency maintenance, dense text rendering, precise editing and aesthetic quality?

This article is from WeChat official account "APPSO", the author is APPSO who discovers tomorrow's products, 36Kr publishes this article with authorization.