HomeArticle

Just now, Gemini 4 Pro has been launched ahead of schedule, far outperforming Astra and Fable.

新智元2026-09-18 11:07
Gemini 4 Pro has been anonymously leaked ahead of official release, with capabilities that far outperform rivals including GPT-4 and takes a comprehensive lead across all dimensions.

What a well-hidden powerhouse!

This time, Google finally laid all its cards on the table.

Just in the past couple of days, a new model quietly entered the large model arena Arena, with the name marked as "gemini-3.8-flash".

After several rounds of hands-on tests, leading AI experts and developers almost gasped in surprise —

This is very likely Google DeepMind's next-generation flagship Gemini 4 Pro, which has been iterated for a full half year!

A benchmark chart that went viral across the internet can be described as "horrifying".

Gemini 4 Pro is performing overwhelmingly, with its capabilities in coding, agents, reasoning and other fields completely outperforming GPT-6 Astra and Claude Fable 5.1.

After experiencing it, developer Vidhi was completely shocked by Gemini 4 Pro.

An SVG image of a cat proves everything. GPT-6 Astra directly output a "big cat" that looks more like a tiger than a cat.

Just as the rumors say, Google has implemented RSI internally, which allows Gemini 4 Pro to compete head-on with Astra and Fable.

The long-dormant behemoth has truly made a strong comeback!

Gemini 4 Pro runs anonymously, outperforms Astra and Fable

The AI circle hasn't been this excited for a long time.

It is worth noting that Google's last major version update, Gemini 3.1 Pro, was released 7 months ago on February 19.

At the Google I/O conference, Sundar Pichai once announced that Gemini 3.5 Pro would be launched in June.

No one expected that this major release was postponed repeatedly, and ultimately got scrapped before launch!

Earlier this month, the Wall Street Journal exclusively broke the news that Gemini 3.5 Pro was cancelled for a very simple reason —

The evolutionary progress of 3.5 Pro is not even as good as that of the Flash model.

Fortunately, Gemini 4 has achieved excellent evaluation in pre-training, and post-training is proceeding smoothly.

Now, Gemini 4 Pro has put on the "disguise" of gemini-3.8-flash, and its first checkpoint has appeared on Arena.

After all, the Gemini 3.8 Flash model was released as early as the beginning of September, so the reappearance of the same name is extremely unusual.

A large number of developers after hands-on experience are deeply impressed by 4 Pro, and even feel that its performance surpasses Astra and Fable.

From the leaked benchmark test chart, it can be found that 4 Pro has achieved comprehensive lead with a very impressive report card —

DeepSWE v1.1: On the AI agent coding task, 4 Pro scored a maximum of 88%, nearly 2% higher than Astra;

GDPval-AA v2: In real knowledge work tasks, 4 Pro is the only model that obtained 2064 Elo;

Terminal-bench 2.1: 4 Pro also has the strongest terminal coding capability at 95.3%;

OSWorld-2.0: In terms of computer usage capability, 4 Pro scored 86.8%, beating Astra and Fable.

If the scores on this table are true, Gemini 4 Pro is truly the "strongest AI on the planet".

In terms of price, compared with Astra and Fable, it is also the most cost-effective one among the top three AI models. The price is 2.25 dollars per million input tokens, and 11.25 dollars per million output tokens.

After accessing the backend of Gemini 4 Pro, AI researcher Qwinah found that it has a 10 million input upper limit, a 256 thousand output upper limit, permanent cross-session memory, and can connect to the internet without an API.

First public test of 4 Pro across the network hits hot search

What is more notable than the benchmark scores is a large wave of amazing practical tests across the internet.

UI web design, taste greatly improved

After testing, developer Bee even exclaimed "Gemini 4 Pro is too unreal......"

In just 14 minutes, 4 Pro created a "creative display webpage" with the texture of sketch, pencil and graphite as the theme.

It can even conceptualize the idea of "scrolling as brush strokes" on the webpage: when you scroll down, the lines gradually darken, bringing extremely smooth interactive experience.

The following is a work website that combines cyberpunk style and retro futurism.

In the design process, 4 Pro combines 3D grids, data dashboards and interactive 3D models on the first screen.

3D + SVG performance is extremely outstanding

In the classic viral test of "pelican riding a bicycle" spread across the internet, the output effect of Gemini 4 Pro is very shocking.

The color matching, day and night switching, car lights, anatomical annotation and cadence control all have a completion level that is no less than that of top large models.

After experiencing it, Pankaj Kumar directly summarized several improvements:

It runs very fast, SVG and 3D generation have made obvious progress; it can accurately understand complex requirements with only one prompt.

Even, it can now directly generate small games with complete interactive logic and good visual effects.

Take another look at the pixel-style 3D pagoda test, which brings an immersive feeling.

The following is the 3D model of Airbus H145 helicopter made by 4 Pro in 10 minutes.

In addition, in a 3D flight simulation test, the picture effect of 4 Pro greatly surpasses Gemini 3.8 Flash.

This alone can prove that the official Gemini 3.8 Flash is not the same model as the gemini-3.8-flash on the Arena.

Output playable games with one click

Gemini 4 Pro also performs very well in game generation.

From page construction to the design of each module in the large game, the "Minecraft" generated by it is no less than the previous test result of Astra.

There is also the following 3D kart racing game, which looks more modern than the once-popular racing game KartRider; and this work was made by Gemini 4 Pro in a very short time.

The core technology Google really bets on is RSI

Once the RSI cycle can operate continuously, the rhythm of model iteration may be completely changed.

Last month, Jasjeet Sekhon, Chief Strategy Officer of Google DeepMind, made it clear directly at an event at UC Berkeley —

RSI (Recursive Self-Improvement) has become a key part of the huge investment logic in the AI industry.

The whole industry is betting that the capability curve will continue to rise.

If recursive self-improvement is truly achieved, this curve may even accelerate further from "exponential growth".

Recently, more radical statements are spreading across the internet:

The reason why Gemini 4 completed pre-training ahead of schedule is that Google DeepMind has achieved "RSI closed loop" during training.

Interestingly, on the 14th, a Google research named Dream-RSI was also made public.

It studies how to make Agent continuously improve its search strategy during the exploration process, and upgrade the next round of exploration from experience.

It can be confirmed that Google DeepMind has publicly placed RSI in a very important position.

Moreover, there are more and more signs that AI is participating in improving AI itself.

Are the top three tech giants making a comeback?

Therefore, Gemini 4 Pro is facing a completely changed competitive landscape.

OpenAI has Astra; Anthropic has Fable 5.1.

The two giants have extended the competition from simple chat scenarios to long tasks, Agent, coding, reasoning and even independent scientific research.

It is no longer enough for Google to just make a model "stronger than Gemini 3.8".

Gemini 4 Pro must re-prove one thing: Gemini is still at the cutting edge of the industry!

If Google wants to maintain