HomeArticle

Gemini 3.7 Flash drops unexpectedly, which is the first major move of Google AI after the urgent leadership reshuffle, and the truth of its internal strife has now come to light.

极客邦科技InfoQ2026-08-14 14:38
You read that right: Gemini 3.6, which was released just three weeks ago, has launched a brand new version again.

Google is updating Gemini at an unprecedented pace.

Last night, local time on August 13, Google DeepMind officially released Gemini 3.7 Flash. Google defines it very directly: "the smartest workhorse model to date", with a core focus on coding and Agent scenarios.

Interestingly, the release of 3.7 Flash came only three weeks after the launch of Gemini 3.6 Flash. Google itself specifically emphasized this point at the beginning of its official blog: Gemini 3.7 Flash was launched just three weeks after 3.6 Flash, supported by developer feedback and a series of algorithmic innovations.

More notably, it has been only eight days since Google DeepMind completed its most drastic power restructuring in recent years.

On August 5, Demis Hassabis stepped down as CEO of Google DeepMind and took up the positions of Chairman of Google DeepMind and Chief Scientist of Alphabet; Koray Kavukcuoglu, CTO of Google DeepMind and Chief AI Architect of Alphabet, officially took over the daily operational management.

Strictly speaking, Koray did not get the CEO title, but was promoted to Senior Vice President of Google DeepMind. However, in terms of actual authority, he has become the new "leader" of Google AI: he is fully responsible for Gemini model R&D, Frontier AI research, Gemini App and developer teams, and reports directly to Sundar Pichai. Google later also confirmed to Reuters that Koray will have the final say on major decisions of DeepMind.

Thus, a timeline that is hard to ignore has emerged:

On August 5, DeepMind changed its top leader, and then on August 13, Gemini 3.7 Flash was released.

If we connect the events that happened at Google in the past few months, this Flash update is far more than an ordinary model iteration.

Iteration in three weeks, Google starts to race against time

Google has almost placed all the core priorities of Gemini 3.7 Flash on two keywords: Coding and Agent.

This focus is not surprising at all. Over the past year, the focus of large model competition has rapidly shifted from chat and knowledge Q&A to code generation, tool calling, Computer Use and long-chain Agent tasks. A single win in Benchmark is no longer enough to make users pay for the model, but a model that can call tools, modify codes, read documents, access systems and complete actual tasks for dozens of consecutive steps is still worth expecting.

Gemini 3.7 Flash is exactly enhanced in this direction.

According to Google's official data, in FrontierCode 1.1 Main which measures production-level code quality, the score of Gemini 3.7 Flash reaches 43.6%, a significant improvement compared with 34.4% of Gemini 3.6 Flash.

In DeepSWE v1.1 which evaluates long-cycle software engineering tasks, the performance increased from about 49% of 3.6 Flash to 65.3%.

Web development is also the core focus of this upgrade.

Google states that 3.7 Flash can generate more complete and functional web applications with fewer prompts, and its ability to restore screenshots, images and complete Design Systems has also been improved.

On WebDev Arena, its Elo score increased from 1538 of 3.6 Flash to 1588.

The improvement of Agent capability is even more obvious.

In the Agentic Terminal Coding test of Terminal-bench 2.1, Gemini 3.7 Flash reaches 85.8%, higher than 78.0% of 3.6 Flash; in the more difficult Terminal-bench 3.0, the score jumped from 5.4% to 14.9% in one go.

In AutomationBench which tests the automation capability of real enterprise workflows, the score of 3.7 Flash increased from 17.0% to 30.4%.

Computer Use also saw significant progress. On OSWorld 2.0, 3.7 Flash scored 47.9%, while the score of 3.6 Flash was 33.8%.

In other words, 3.7 Flash does not simply improve the scores of mathematics or knowledge Q&A, but focuses on the areas that Google most needs to strengthen at present: coding, tool calling and Agent execution.

Google even specifically emphasized in its blog that 3.7 Flash will take the initiative to adjust strategies when encountering execution obstacles, clarify user intentions when necessary, and follow instructions more strictly.

The model will "invest more thinking" in multi-step planning and tool calling, reducing manual supervision and repeated retries.

The goal behind this statement is very realistic: the model can no longer only stay at "giving a good answer for the first time", and now it needs to complete the task thoroughly.

Flash is approaching the flagship model

If you only look at the model card released by Google, another change of 3.7 Flash is also obvious: the performance gap between Flash and high-priced flagship models is narrowing.

On the Artificial Analysis Intelligence Index, Gemini 3.7 Flash scored 56, Claude Sonnet 5 scored 55, and GPT-5.6 Terra scored 57.

In FrontierCode 1.1, the 43.6% score of 3.7 Flash is higher than 42.7% of Claude Sonnet 5 and 41.3% of GPT-5.6 Terra.

In the Web Development Code Arena, the Elo score of 3.7 Flash reaches 1588, which is also higher than Claude Sonnet 5, GPT-5.6 Terra and Muse Spark 1.2 listed in Google's model card.

Of course, this does not mean that Gemini 3.7 Flash has comprehensively surpassed the flagship models.

For example, in DeepSWE v1.1, GPT-5.6 Terra still reaches 69.6%; in Terminal-bench 3.0, GPT-5.6 Terra scores 20.8%, which is significantly higher than 14.9% of 3.7 Flash; in the GDPVal-AA v2 knowledge work test, 3.7 Flash also lags behind several competing models.

But this just illustrates Google's current positioning of Flash: it does not need to be the absolute first in all Benchmarks. Instead, it hopes to use capabilities close to cutting-edge models, lower prices and higher throughput to become a model that can run Agents on a large scale.

This is also the truly noteworthy part of the expression "workhorse model".

Google is also starting the Agent cost war

The more aggressive part of Gemini 3.7 Flash is actually its price. Before the end of this year, its introductory price is: $0.75 per million input Tokens, $3.75 per million output Tokens. According to Google, this is equivalent to half the original price of Gemini 3.6 Flash.

This is not a permanent price, but at this stage Google still wants to participate in this cost war.

The Google DeepMind model card shows that the promotional price of 3.7 Flash will end on December 31, 2026.

From January 1, 2027, the price will return to $1.5 per million input Tokens and $7.5 per million output Tokens.

Even so, this pricing strategy still shows that Google is competing more and more seriously for the actual workload of Agents. Because after entering the Agent era, the cost structure of large models has changed.

A chatbot may only call the model once or twice, but a truly working Agent may need to plan tasks, search for information, read more than a dozen files, call multiple tools, retry after failure, and continuously send the execution results back to the model.

Once Agents are scaled up, the number of Tokens and tool calls will expand rapidly.

The party that can run a sufficiently intelligent model for dozens or even hundreds of steps at the lowest cost is likely to be the key to stand out from the competition, and Flash is taking on this role.

On the day Gemini 3.7 Flash was released, Google deployed it to Gemini Spark.

Spark is a personal AI Agent launched by Google at this year's I/O, targeting Google AI Pro and Ultra users, and currently covers more than 160 countries and regions. Google defines it as a personal Agent that can run continuously 24 hours a day and take actions on behalf of users under user control.

After 3.7 Flash goes online, Spark will use the new model to process tool calls such as Google Workspace. The scenarios given by Google include integrating multiple files, drafting emails and updating project status documents.

In other words, Gemini 3.7 Flash is not a model only used to brush Benchmark scores.

On the day of release, it was integrated into Google's own Agent products. At the same time, developers can use 3.7 Flash through Gemini API, Google AI Studio, Android Studio and Google Antigravity; enterprise users can call it through products such as Gemini Enterprise Agent Platform.

But here comes another question: Flash has been updated from 3.6 to 3.7, where is Gemini 3.5 Pro?

The answer is: it has not been officially released yet.

As early as when Gemini 3.6 Flash was released on July 21, Google stated that Gemini 3.5 Pro is being tested with partners and will be made public as soon as it is ready.

At that time, Google even announced that Gemini 4 was undergoing the "most ambitious" pre-training in the company's history.

But until the release of 3.7 Flash, Google still did not announce the official release time of 3.5 Pro.

Reuters also specifically pointed out in its report on August 13 that this release did not provide any new information about when the flagship Pro model will be launched.

This makes the sudden appearance of Gemini 3.7 Flash even more thought-provoking. Because just one day earlier, Reuters had just disclosed more details behind the major restructuring of Google DeepMind.

The truth of the "internal conflict" comes to light

If you have to describe this huge shake-up of Google AI as "internal conflict", it is actually not simply a personal feud between executives, but a simultaneous outbreak of several groups of long-standing contradictions: scientific research and commercialization, London and Silicon Valley, different technical leaders of Gemini, and where the limited computing resources should be invested.

Citing multiple insiders, Reuters reported that in April this year, when Anthropic was advancing the new Claude model, Sergey Brin, co-founder of Google, once asked DeepMind to "speed up" and devote more energy to Gemini in an internal meeting with hundreds of participants.

By August, the new generation flagship Gemini originally planned to be launched by Google had been delayed for about two months.

Insiders said that internal tests showed that it still lags behind some competitors in key capabilities such as coding.

What's more troublesome is the organizational structure.

Reuters said that the Gemini project has long had differences of opinion among multiple leaders, coupled with limited computing resources, some directions that later proved to be very important - including Coding - did not get sufficient resources in time.

At the same time, there is also tension between Google Cloud and DeepMind over the allocation of limited TPU computing resources. Some Google Cloud executives believe that after Koray takes over DeepMind, this resource allocation conflict is expected to be alleviated.

Power is also shifting further from London to Mountain View.

Koray moved from London to Google's headquarters last year and became Alphabet's Chief AI Architect. Some London team leaders then felt that their influence on the Gemini technical roadmap declined.

Compared with Hassabis, Koray has always been considered closer to the product and commercialization system within Google. He used to be the main interface between DeepMind and Google Cloud.

After the restructuring on August 5, this power transfer was formally finalized. Hassabis turned to AGI, scientific research and long-term strategy, while Koray took direct charge of Gemini. Some non-technical teams are also starting to transfer from DeepMind to the Google corporate system.

On X, a netizen appropriately described DeepMind's independence as a kind of "privilege" in peacetime. As soon as the market shows any signs of unrest, this "privilege" is in jeopardy. He wrote:

"DeepMind's ability to remain independent is, in the final analysis, just a 'privilege' in peaceful days. As soon as Alphabet panics, the 2014 agreement is directly voided, and Brin personally steps in to hold an all-hands meeting."

At the same time, many of the most important figures in Google AI chose to leave. Jeff Dean, Sanjay Ghemawat, Oriol Vinyals and Quoc Le left to create Discovery Loop, among whom Jeff Dean and Oriol Vinyals were once the original co-technical leaders of Gemini.

Earlier, Noam Shazeer had joined OpenAI, and John Jumper, the core researcher of AlphaFold, joined Anthropic