Google launched a surprise attack on OpenAI late at night, outcompeting its Pro service at a rock-bottom price.
Just now, Google dropped Nano Banana 2.1 deep in the night!
A Flash-class small model that even outperforms its own flagship Nano Banana Pro.
Key highlights:
Image editing: Supports mask-based local editing, modify only the selected area while keeping all other parts completely unchanged
Consistency: Characters and products remain unchanged across multiple rounds of edits, supporting up to 14 reference images at maximum
Image quality: Generated visuals are far more photorealistic, with native 4K output available at highest
Pricing: The image generation cost is half of that of Nano Banana 2, at only $0.034 per 1K resolution image
Benchmark ranking: It takes 4th place in the Arena multi-image editing leaderboard, with only three GPT Image models from OpenAI ahead of it
Rollout: Gemini App, Search AI mode, AI Studio, Flow and other platforms will support this new model successively starting today
Flash outperforms Pro, is Nano Banana positioning for cost-effectiveness?
First of all, Nano Banana 2.1 is still a member of the Gemini 3 model family, built on Gemini 3.6 Flash, with a knowledge cutoff date of March 2026.
According to the previous naming convention, Nano Banana 2.1 can also be named Gemini 3.6 Flash Image
Based on official benchmark tests, Nano Banana 2.1 has taken a clear leading position in visual design, editing, consistency and other dimensions.
Although there is still room for improvement in small text rendering, 3D reasoning and factual accuracy, this is undoubtedly a very impressive update for users who need image generation and editing services:
Create and edit images with professional-level precision and control, and support multiple fast iterations
Generate clear text for posters and complex charts
Long-context real-world knowledge
Localized text rendering across multiple languages
The blind voting results from netizens also prove this point.
On the Arena leaderboard, Nano Banana 2.1 ranks 4th in multi-image editing, 5th in text-to-image generation, and 6th in single-image editing.
Especially on the multi-image editing leaderboard, only three GPT Image models from OpenAI are placed ahead of it.
In comparison, the previous generation Nano Banana 2 only ranked 7th, 11th and 14th on these three leaderboards respectively.
Swipe left and right to view
Considering that its image generation price is only half of that of Nano Banana 2, it seems that Google is ready to launch a price war.
Modify only the selected area, no character face distortion at all
Mask editing is a familiar feature for all Photoshop users. You select a specific area, and only that area will be modified, while the rest of the image stays as it is.
But this has long been a pain point in AI image generation.
When you tell the model "replace the cup in the lower left corner with a red one", it first has to guess which cup you are referring to, then regenerate the whole image, and whether the other parts can be preserved entirely depends on luck.
Now you only need to circle the target area. Changing clothes, replacing store signs, erasing passersby in the background, all operations only affect the selected part.
Netizen ibexdream gave it three consecutive tasks: generation, image editing, and style conversion.
The first task: draw a banknote, on which there is an old philosopher resting his chin in thought, with his hair extending backwards to form a whole maze.
The second task is to edit on this existing image: change the serial number from "№ 1098471" to "NB-21098471", replace the small text on both sides, add a banana as a phone in the old philosopher's hand, and add a ribbon printed with "IN BANANA WE TRUST" at the bottom.
Compare the two images, the maze, robe and beard all remain completely unchanged.
The third task is style conversion, turning the whole banknote into an oil painting.
In the model card scoring, the mask editing item of version 2.1 gets 1049 points, 84 points higher than Nano Banana 2, and 122 points higher than Pro.
Even without enabling the reasoning mode, 2.1 still outperforms Pro in every single dimension.
Another major upgrade is subject consistency, which essentially means no unexpected face distortion.
People who shoot e-commerce product photos, make brand posters or draw serialized comics must have deep experience: when the same model changes into ten different sets of clothes, the face will start to look abnormal from the third outfit.
This item sees the most significant score increase this time. The multi-person consistency score reaches 1106, 128 points higher than Nano Banana 2.
Developers can input up to 10 object reference images and 4 character reference images for the model to follow when generating.
Netizen BG-VC tried this feature: he provided a portrait photo of the model, plus 5 reference images of separate items including top, skirt, bag, shoes and earrings, and asked the model to generate 6 photos in different scenes, all 6 images were generated in one go.
Walking, sitting, leaning against the table edge, with six different poses, the face, clothes, bag and earrings all remain exactly the same, the generated set of photos is fully qualified to be used as official e-commerce product visuals.
The generated images themselves also look far more like real shots. In the official sample images, the beetle covered with dewdrops and the girl in green dress on the stairs outside the pink house are almost impossible to be identified as AI-generated at first glance.
In addition, when generating infographics, it can first look up materials via Google Web Search and Google Image Search before starting to draw.
In Google's automatic evaluation, the factual accuracy score of version 2.1 for infographic generation is 0.521, while Nano Banana 2 only gets 0.179, and Pro gets only 0.265.
In other words, when you ask it to generate popular science graphics, the number of content errors will be greatly reduced.
Image generation gets cheaper, reasoning capability gets more expensive
Image generation is priced at $30 per million tokens, exactly half of the price of Nano Banana 2.
The cost per 1K resolution image drops from $0.067 to $0.034, and the cost per 4K resolution image drops from $0.151 to $0.076.
This price is exactly the same as the low-priced Nano Banana 2 Lite released in July.
However, the pricing of other items has increased instead.
The input price rises from $0.5 per million tokens to $1.5, which is 3 times the original price. The output price for text and reasoning also rises from $3 to $7.5.
It is equivalent to Google shifting its revenue focus from "image generation" to "reasoning capability".
Currently, developers can select this model in AI Studio right now, the model name is gemini-nano-banana-2.1.
The API of the previous Nano Banana 2 will be shut down on October 29.
Ordinary users don't need to wait either.
Gemini App, Search AI mode, Flow, design tool Stitch,