Gemini 4 Pro is suspected to have been leaked, and the claim of "AI deceleration" has once again become empty talk.
In the past few months, OpenAI and Anthropic have successively pushed their flagship models forward one after another, while Google has remained unusually quiet.
Flash has been updated roughly every three weeks, with successive iterations of 3.6, 3.7, and 3.8 rolling out, but the Pro variant that represents the highest capability ceiling has not shown any signs of progress for a long time.
It was not until the past couple of days that a model labeled with the name gemini-3.8-flash suddenly popped up in Arena, the large model blind test arena.
As soon as developers got their hands on it, they noticed something odd. The real Gemini 3.8 Flash has already been released, and everyone knows exactly what performance level it should reach. But this "3.8 Flash" delivered a noticeably higher performance in writing code, generating SVG, and running Agent tasks.
After several rounds of actual testing, a speculation quickly spread across the community:
This model is very likely the next-generation flagship Google has hidden behind the label — Gemini 4 Pro.
Shortly after that, a more exaggerated benchmark comparison chart went viral online. It outperformed GPT-6 Astra and Claude Fable in multiple metrics including coding, Agent operation, and reasoning.
Just a few days ago, several top AI companies were still publicly discussing whether they should slow down the pace of development. Now, before the calls for deceleration even land, a new round of competition has already sped up at full throttle.
Behind this "god-tier model" suspected to be Gemini 4 Pro, there is another more dangerous keyword:
RSI.
Suspected Leak of Gemini 4 Pro
On September 2, Google officially released Gemini 3.8 Flash. The official statement says this is the latest Flash variant in the Gemini 3 series, with significant improvements in software engineering, Agent tasks and complex multi-step reasoning; only three weeks have passed between the release of 3.7 Flash and 3.8 Flash.
Therefore, when another model with the exact name gemini-3.8-flash appeared in Arena, developers quickly noticed the anomaly.
The official 3.8 Flash is already publicly available as a ready reference for everyone. This new model is obviously far more powerful, with capabilities that clearly belong to the Pro model line.
The first batch of actual tests that sparked widespread discussion focused on SVG generation, web page design and 3D scene creation.
Developer Harshith asked the model to draw a side view of a domestic cat with SVG, and compared the result with outputs from GPT-6 Astra Max and the official Gemini 3.8 Flash. The version in Arena showed visible gaps with the official 3.8 Flash in terms of outline, proportion and detail completeness.
From left to right: "4 Pro", Astra, 3.8 Flash
X user @thtbee_ conducted a series of tests on this suspected Gemini 4 Pro model from multiple perspectives.
For a Voxel-style pagoda, the model ran for 8 minutes and directly generated a full set of interactive 3D scenes.
For an Airbus H145 helicopter, it only took about 10 minutes to build a complete 3D display page with extremely rich details. However, the user noted that the UI elements at the bottom of the page "looked a bit rough".
In web design tasks, the model spent about 14 minutes making a monochrome theme website that is very clean and complete overall. The long-criticized poor design taste of Gemini seems to have finally been fixed this time.
Another user @Mr_Salio directly used this suspected Gemini 4 Pro model to develop games. The finished product runs smoothly, with high completion of world generation and game design.
A more critical clue came from developer Qwinah.
He claimed that he successfully "ghost routed" to the backend of Gemini 4 Pro, and posted a screenshot of suspected internal terminal information.
According to his statement, the internal identifier of this model directly shows G4P-ARGON and VIA_3.8_FLASH, with a maximum output of 256K, a context window of more than 10 million tokens, and capabilities including cross-session permanent memory, offline web access without API, automatic backend script compilation and execution, and even physical robot control.
If this information is true, it is obviously not an ordinary Flash version upgrade.
At the same time, a benchmark comparison table of Gemini 4 Pro also began to go viral in the community.
According to the data in the chart, Gemini 4 Pro scored 88.7% in the software engineering test DeepSWE v1.1, reached 95.3% in the terminal Agent test Terminal-Bench 2.1, hit 72.1% in the expert knowledge reasoning test HLE-Verified, and achieved 86.8% in the computer operation Agent test OSWorld 2.0.
More exaggeratedly, on the GDPval-AA v2 metric that measures real knowledge work performance, it is marked as 2064 Elo, directly breaking through the 2000 point mark.
If these data are true, Gemini 4 Pro can easily outperform Fable and Astra, and directly stand at the top of cutting-edge large models.
But the more exaggerated the claim is, the more cautious we need to be.
It is worth noting that this benchmark table does not fully match the suspected backend screenshot "dug out" by Qwinah; the score of Astra as a reference item in the benchmark table also does not conform to the official data; neither of the two materials has official endorsement from Google, Arena or relevant benchmark institutions, and they are still only information circulating in the community for now.
The only confirmed fact is that there exists a model named gemini-3.8-flash whose performance completely does not match the official Gemini 3.8 Flash.
There is also a very important reason why people quickly associate this anomaly with Gemini 4 Pro: Google's Pro model line has been vacant for far too long.
Gemini 3.5 Pro has not been officially launched for a long time, and Gemini 4 has long been confirmed to enter the training phase. In the past few months, the Flash variants have been iterated generation after generation, but no new model in the Pro line that represents the upper limit of capability has ever appeared.
Now, an abnormally high-performance "3.8 Flash" suddenly emerged from Arena.
As a result, the speculation became more and more reasonable: Google may not have planned to leave the major upgrade to 3.5 Pro at all, but directly saved it for Gemini 4 Pro.
Behind Gemini 4 Pro,
The Bigger Keyword Is RSI
If the existence of Gemini 4 Pro is still just a community rumor for now, what is more noteworthy is that Google is significantly accelerating the speed of AI participating in model R&D.
In all the recent rumors about Gemini 4 Pro, the keyword RSI has been mentioned repeatedly.
RSI is the abbreviation of Recursive Self-Improvement. Simply put, it means that AI begins to participate in the creation of stronger AI, and the more powerful AI further accelerates the R&D progress of the next-generation model.
Once this cycle starts running, it is not just the model capability itself that gets accelerated.
Even the speed of model evolution will be accelerated by the model itself.
Reuters disclosed in August this year that Google co-founder Sergey Brin has been pushing for the acceleration of Gemini development in recent months, and listed recursive self-improvement as one of his key focus areas.
Although Google has not yet publicly announced that it has achieved this closed loop, it is handing over more and more work that originally belonged to researchers to Agents, including model evaluation, finding improvement directions, running experiments, and feeding the results back to the model R&D process.
When releasing Gemini 3.8 Flash on September 2, Google also mentioned in the official note that a long-running Agent loop is "recursively evaluating and improving the underlying model".