HomeArticle

The global showdown over foundational models is in full swing, and the AI community will not get to enjoy the week-long National Day holiday this year.

量子位2026-09-20 15:10
Claude, Gemini and Kimi are all gearing up to roll out their blockbuster new moves.

We're already giving you a heads up — the 13-day holiday spanning Mid-Autumn Festival and National Day is likely to be the busiest, most explosive two weeks in the global AI industry this year.

Friends, skip the vacation, skip the sleep, grab your comfiest chair and get settled in.

"Foundation Model Showdown Week" is truly coming!!

The undercurrents are already surging in Silicon Valley.

News of the gray test for Anthropic's new Claude has long been spreading far and wide.

Google is even more straightforward: the suspected Gemini 4 Pro has secretly snuck into the Arena under a fake name to take placement tests.

On the domestic front, everyone is holding back their biggest moves right before the holiday.

Kimi has been dropping cryptic hints about "K3.1", DeepSeek has already written its next Pro card into official documents in advance, and Qwen4 has also directly unveiled its new architecture.

Right in the center of the ring, OpenAI just released GPT-6 Astra earlier this month — a foundation model that made OpenAI pop the champagne for AGI ahead of schedule.

Wow, this whole setup looks like everyone made an appointment in advance:

All the models that are supposed to arrive, that might arrive, and that have already snuck into the examination room secretly, are gathering at the same time.

The last big show before the holiday, Are you ready?

Silicon Valley teams are ganging up on OpenAI GPT-6 Astra

Over in Silicon Valley, Anthropic, Google, and Elon Musk's Grok have all reported new upcoming moves.

Anthropic's Claude 5.2 is on the verge of launch

Unsurprisingly, Anthropic is the first one that can't hold back.

The latest exclusive report from Reuters has confirmed that Anthropic is considering launching a new model, though the exact release date and official name have not been made public yet.

The most widely expected models in the community are Claude Opus 5.2 and Fable 5.2.

The reason these two models have drawn the community's attention is a series of recent gray test signals that have emerged.

The earliest clue came from Claude Code.

Some developers found that even though they had selected Opus 5, its performance in some complex tasks was noticeably different from previous versions.

Then the community started to speculate:

Has Anthropic been quietly switching models in the background, routing part of the requests to the new internal checkpoint?

This speculation was soon confirmed by more clues.

Multiple developers reported that requests originally calling Fable 5.1 were silently redirected to the new version in the background.

Some people tested it in high inference mode and found that the output quality of the new model is significantly better than GPT-6 Astra, at the cost of slower speed and higher cost.

The most dramatic part is the "midnight surprise" of Opus 5.2.

On the evening of September 17, Anthropic suddenly cut off the gray test channel for Opus-Next (rumored to be the internal code name for Opus 5.2), the number of active test accounts dropped to zero for a while, and the developer community was full of complaints.

But only a day later, the gray test went back online, and the scope expanded from Claude Code to Chat and Cowork.

The big background behind this series of actions is:

GPT-6 Astra released by OpenAI earlier this month has already captured about 13% of enterprise AI spending, while Anthropic's Claude Fable accounts for about 8%.

Anthropic choosing to spread the news of the new model on the eve of its IPO is clearly an attempt to seize back the narrative initiative.

Google Gemini 4 Pro internal test results are going viral

Google's move is even more full of dramatic effects.

A model named gemini-3.8-flash recently appeared in Arena.

The name looks unremarkable, but many testers soon noticed something wrong:

This thing is not like Flash at all.

Everyone unanimously inferred that it is very likely the hidden checkpoint of Google's next-generation flagship Gemini 4 Pro, and the community also spreads that its internal code name is Argon.

A bunch of demos that emerged afterwards directly pushed the discussion heat to a new high.

For the classic AI industry test question "Pelican Riding a Bicycle", the suspected Gemini 4 Pro version not only draws the bird and the bicycle, but also further builds a complete interactive page with pedaling movements, lighting changes and control components.

Compared with the regular 3.8 Flash, the complexity is visibly much higher.

Another example is Drawing PS5 SVG: the model spent about 10 minutes directly drawing a set of extremely detailed PS5 vector graphics with code, including complex curves, shadows and body structure.

Testers directly commented that this is one of the most exaggerated outputs they have ever got from an AI model.

There is also the Voxel Pagoda, this time the model directly built a complete 3D scene:

The multi-layer pagoda, terrain, trees and surrounding environment are all generated together, and the functions of view adjustment and lighting control are also retained.

Compared with previous checkpoints, the geometric structure and the completeness of the entire scene are significantly improved.

And a widely spread demo Side View Domestic Cat SVG, the prompt is very simple:

Draw a domestic cat in side view.

Testers put the results of the suspected Gemini 4 Pro side by side with GPT-6 Astra Max and the official version of Gemini 3.8 Flash, mainly comparing the outline, limb proportion and spatial relationship (the following pictures are arranged in this order).

The result also got hundreds of likes on Reddit, becoming one of the most popular comparison charts in this wave.

After that, the tests have evolved from SVG to web pages and 3D scenes.

Some netizens even used it to make a small game "Ultraman Riding Godzilla Fighting Monsters":

Built with SVG+HTML, it fires lasers wherever the mouse clicks, and can also accelerate.

Of course, these demos only prove that this mysterious checkpoint in Arena is indeed very powerful.

Whether it is Gemini 4 Pro or not remains to be confirmed before Google's official announcement, but at least judging from the test density —

The next generation of Gemini is already very close.

Elon Musk's Grok 4.7 is being cooked

Elon Musk is not idle either.

Earlier, a netizen asked about the progress of Grok 4.7, and Musk directly replied:

Grok 4.7 needs a few more days to cook.

Counting the days, it will be ready in just a few days.

But according to Musk's previous style, he usually waits for everyone else to release their products before showing up, playing the role of "the mantis stalks the cicada, unaware of the oriole behind" (doge).

Place your bets, I'll put my money on it first.

The domestic market is equally bustling

On the domestic front, some players have already taken the lead —

Just now, StepFun suddenly released its new flagship model Step 5 Preview.

The model adopts a sparse MoE architecture, with a total parameter count of 600B and 27B activated parameters, supports a million-token context window, and is optimized for real-world tasks such as programming, Agent development, and professional knowledge work.

In the globally authoritative Artificial Analysis Intelligence Index, Step 5 Preview scored 44 points, ranking among the top three open-source models in the world, and the cost per task is only 1/8 of Claude Opus 5.

Right after StepFun laid its cards on the table, the answer for Kimi is almost out in the open.

Recently, its official account suddenly released a string of numbers: 415926535897932384626433832795……

At first glance it looks like a random string, but it makes perfect sense when you compare it with the value of pi π.

The beginning of π is:

3.1415926535897932384626433832795……

Remove the leading "3.1", and the rest is exactly this string of numbers.

K3.1? Alright alright, who has already got the hint~ (But some netizens bet the other way, saying this means there will be no 3.1 release)

Apart from Kimi, DeepSeek has also hidden easter eggs long ago.

When DeepSeek V4.1 Flash was released on September 10, the official document clearly mentioned that:

Before V4.1 Pro goes online, part of V4 Pro requests will be routed to V4.1 Flash.

In other words, the existence of V4.1 Pro is basically confirmed, and we are just waiting for its official release.

Judging from its previous release rhythm, V4.1 Pro is most likely to be launched in the window from the end of September to the beginning of October.

There is usually not a long interval between Flash and Pro versions. When the V4 series was released earlier, the Flash and Pro versions were launched almost simultaneously.

In earlier version iterations, after the lightweight version was released first, the flagship version would also follow within a few weeks.

Add the holiday factor to it, and everyone is no stranger to the routine of "releasing new models right when the holiday starts".

Next up is Alibaba's Qwen.

When Alibaba open-sourced Qwen3.8-Flash-Next at the end of August, it