Breaking: Google Earth has urgently withdrawn the Nano Banana 2 image generation feature.
The cyber version of "instantly turning words into reality" has reached a brand-new realm!
Yesterday, Google officially integrated its latest image generation model Nano Banana 2 into the web version of Google Earth.
Who could have expected that when people woke up the next day, the new feature was overused and broken by netizens, and Google urgently rolled it back!
As the generated effects were far too realistic, after strong objections from experts, Google urgently withdrew the new feature and re-released it after "enhancing protective measures".
New gameplay of Nano Banana 2: A visual carnival for netizens
You open Google Earth in your browser, and skillfully move the perspective down to the old street where you played in your childhood. You don't open complicated Photoshop, nor do you copy the screenshot to other AI chat tools.
You just click the brand new, glimmering "Create image" button in the upper right corner.
Let's first look at several official examples given by Google.
Turn back time to restore history
It helps students no longer just read about history, but truly "see" the past.
Teachers can show a historical site in class, then input: "Generate a photorealistic picture that shows the original look of the Pompeii ruins in 78 AD."
The modern relics will instantly turn into a bustling, colorful ancient Roman street scene, allowing students to intuitively see the world 2,000 years ago.
Virtual tour guide: Gain in-depth understanding while exploring
Without leaving Google Earth, you can create customized infographics to learn more about a place.
Input: "Create an easy-to-understand infographic of the Statue of Liberty, and mark key historical facts."
In the background, Gemini will retrieve relevant historical information, and Nano Banana will generate the corresponding visual chart in a few seconds.
"Dream City Renovator": Create professional-grade real estate planning schemes
When clients can intuitively see the final effect, the presentation of real estate projects will become much simpler.
Architects and urban planners can now present the design scheme directly with one sentence: "Reimagine this vacant land in Tokyo as a vibrant shopping and business district with open public spaces."
The originally barren concrete vacant land will be replaced by a high-quality 3D-rendered modern urban oasis, helping clients see the possibilities in the future.
Visualize project results in advance before construction starts
Whether it is a dream backyard studio or a lakeside residence, it is often hard to imagine what it will look like after completion.
Now you can quickly turn it into a preview of the real result.
Zoom in on a vacant land, then input: "Add a modern lakeside cabin built with local sustainable materials."
The system will generate a photorealistic render that perfectly integrates the future residence into the real terrain environment.
Give your favorite place a future renovation
Sometimes, letting your imagination run wild is a pleasure in itself.
If you want to see what a place will look like in 100 years, you can input: "Transform the Google Mountain View campus into a futuristic sci-fi utopia with glowing walkways, glass ecological domes, flying transport pods, and trees growing around the buildings."
Shortly afterwards, the entire area will gradually turn into a cinematic cyberpunk futuristic city.
Users exclaim "Oh My God"
David Gewirtz, an editor of a tech media, once spent hours exploring on Google Earth, indulging in the ability to travel around the world without leaving home. The Street View function, though somewhat intrusive, also fascinated him, and it is still the case to this day.
After learning the news, he immediately tested the new feature, which "led to a lot of hilarious results".
The Independence Hall of the United States is a National Historic Landmark of the US and a world cultural heritage. On July 4, 1776, the United States Declaration of Independence was passed here.
Undoubtedly, it has great historical and cultural value for the United States.
But David Gewirtz had a whimsical idea: "Imagine this is a doomsday dystopian future 500 years later."
With just one sentence, the former landmark was directly turned into "doomsday ruins".
One of the most iconic Philadelphia traditions is the oldest annual folk parade in the United States, the Mummers Parade that has been a New Year tradition since 1901.
This made him think: what Philadelphia needs in the future is a doomsday dystopian holiday Mummers Parade.
So he added a prompt: "Now let it be occupied by zombies, evil clowns and giant alien mechs."
Zombies, evil clowns and giant alien mechs, oh my goodness!
Zombies, evil clowns and giant alien mechs are all at daggers drawn, the atmosphere feels totally off. But one more prompt will make it perfect. He told the AI, "Make all of them cheerful."
Now he believes: "This is a believable future Philadelphia, especially on January 1st."
Please note that he also reminded at last: "While this can bring some interesting experiments, it is obviously not a reliable and practical tool yet."
At present, the new feature does not support direct use under Street View mode at all, which greatly reduces its practicality and fun, which is quite disappointing.
Some media even mercilessly teased: For professional developers or architects, these paintings that cannot guarantee the rigor of spatial geometry and are purely pieced together by probability are nothing more than exquisite "AI Slop".
Unexpectedly, the new feature was broken by netizens in less than a day, and Google urgently rolled it back!
Technical decryption
From "creating something out of nothing" to "geospatial constraint"
In the early logic of AI image generation, AI was like playing a game of pixels.
You give it a prompt, and it pieces together a picture from a huge probability distribution. It doesn't care about gravity, let alone whether the building in the picture really exists in reality.
But Nano Banana 2 integrated into Google Earth completely changed the rules of the game. This is a brand new attempt called "Geospatial Grounding".
At the technical bottom layer, it relies on Gemini's powerful multimodal and Search Grounding capabilities:
The first is real-time capture of the physical base map.
The underlying input that Nano Banana 2 receives is not an ordinary top view, but a composite constraint matrix of "satellite/aerial original image of the current viewport + 3D elevation and terrain grid + spatial camera parameters".
It is speculated that its generation task is strictly anchored:
The terrain cannot be tampered with: the folds of the mountains, the baseline of the lake edge, and the exact position and proportion of the building base must strictly conform to the real geographic situation.
The perspective and spatial relationships must match: newly generated buildings, vegetation or scene elements must be embedded into the real 3D terrain and existing structures, instead of floating in the air.
Subject consistency: Nano Banana 2 supports highly consistent rendering of up to 5 main subjects and 14 objects, ensuring that there will be no obvious visual collapse when the perspective of the new city or new community is switched.
This is conditioned generation under real geographic constraints. The trade-off is that it takes up to two minutes for Google Earth to generate each image.
The second is real-time retrieval of the world knowledge base.
Nano Banana 2 is not just an image generator, it is connected to Google's huge search engine and geographic knowledge base in the background. When you try to "restore history" or "mark landmarks", it will actively retrieve the cultural background and geographic common sense of the landmark.
For example, the user requests to generate an "infographic of the Statue of Liberty".
The model not only draws the picture, but also calls Gemini's world knowledge base to automatically fill in content such as year of construction, materials used, building dimensions, and historical information.
However, whether the information retrieval is completely accurate is another issue that needs to be treated with caution.
The fact that AI can generate seemingly professional infographics does not mean that all facts are absolutely reliable.
The third is high-fidelity details and accurate text.
Nano Banana 2 supports high-fidelity image output of 2K or even 4K, and has extremely strong text rendering capabilities. This means you can directly see neatly typeset, garbled-free text signs or information explanations in the generated scene, and even directly "hand-draw" popular science infographics with real historical facts on landmarks.
At the operational level, Google has compressed this complex technology to the extreme: you can view the changes by moving forward/backward during generation, click "Refine image" to add prompts for further iteration, and the interface provides an intuitive before/after experience.
This directly saves ordinary people from the tedious workflow of "screenshot-export-input to other AI-recombine", allowing "creativity" and "real geography" to achieve seamless native integration.
After you are satisfied, click "Save to Project", the generated image will be permanently pinned to the personal project layer in the form of Placemark, retaining the camera perspective. Anyone who clicks into this coordinate can see your "another possibility".
Does the virtual Earth help Google Gemini stand out from the competition?
The significance of Nano Banana being integrated into Google Earth may go far beyond a new feature.
The AI image generator track is in the middle of a scuffle.
OpenAI's GPT-image-2 leads the Arena ranking, and Nano Banana 2 Lite ranks fifth. Adobe Firefly, Midjourney