HomeArticle

Google Gemini 4 has been launched, dropping three new releases in one night.

新智元2026-07-22 08:23
For those of us who have shelled out real money purchasing Tokens and building Agents, there is only one word that is truly practical for today's launch — reduction.

Three back-to-back launches in one go, Google is no longer holding back.

Today, Google DeepMind has dropped three major announcements at once, unveiling three aces in one go:

  • Gemini 3.6 Flash
  • Gemini 3.5 Flash-Lite
  • Gemini 3.5 Flash Cyber

A more powerful main model, a faster lightweight variant, and a dedicated security specialist built exclusively for vulnerability remediation.

These three Gemini models point to the same goal: making AI Agents running in production environments faster, smarter, and more affordable.

And on the very same release day, Google also dropped an even bigger bombshell —

Google DeepMind has "kicked off its most aggressive pre-training campaign in history," targeting the next-generation model Gemini 4.

On one side, three new Gemini models are released in quick succession; on the other, Gemini 4's development is already in full swing.

The signal Google is sending could not be clearer: in this ultimate high-stakes game of AI, the real show has only just begun.

Google's Late-Night Major Reveal

Three Gemini Models Unveiled Together

Let's break down each of these three Gemini models one by one.

Gemini 3.6 Flash: Cuts Token Consumption by 65%

First, let's look at the true "centerpiece" of this launch — Gemini 3.6 Flash.

Its core highlight is drastically reduced token usage.

According to measurements from the Artificial Analysis Index, it uses 17% fewer output tokens compared to the previous generation 3.5 Flash.

On coding benchmarks like DeepSWE, it can save up to 65% of tokens.

Not only does it produce fewer unnecessary outputs, it also requires fewer reasoning steps and tool calls to complete multi-step tasks, taking fewer detours — which translates to both faster speed and lower costs.

On top of that, while Gemini 3.6 Flash cuts token consumption, its price has been reduced even further:

$1.5 per million input tokens and $7.5 per million output tokens, making it cheaper than 3.5 Flash.

Greater efficiency, lower cost, and stronger performance — that's what real sincerity looks like. 3.6 Flash outperforms its predecessor across multiple key benchmarks —

  • DeepSWE: Programming test performance rose from 37% to 49%
  • MLE Bench: Machine learning research test performance increased from 49.7% to 63.9%
  • OSWorld-Verified: Direct computer operation test performance improved from 78.4% to 83%
  • GDPVal-AA v2: Completed 1421 knowledge-intensive work tasks, over 70 points higher than the previous generation

In the demo below, 3.6 Flash demonstrates its exceptional multi-agent orchestration capabilities: it not only easily handles complex code migration tasks, but also delivers a dimensionality-reduction upgrade over 3.5 Flash in both response speed and code quality.

Powered by Gemini Canvas, 3.6 Flash enables a full 3D workflow, allowing users to manually build a professional photography-grade texture extraction tool in minutes.

It can also transform into an "AI interaction designer" to effortlessly create an incredibly immersive themed studio.

3.5 Flash-Lite: Outperforms Its Older Sibling

The second model, Gemini 3.5 Flash-Lite, takes a different path: prioritizing extreme speed and ultra-low cost.

3.5 Flash-Lite delivers an output speed of 350 tokens per second, the fastest in the entire 3.5 series.

Its price is astonishingly low: $0.3 per million input tokens and $2.5 per million output tokens.

Blazing fast and extremely affordable, it is purpose-built for high-volume, high-frequency scenarios like massive document processing and Agent-powered search.

Most impressively, it has managed to outperform its own "older sibling" in the product lineup.

Across multiple programming and Agent benchmarks, this lightweight small model outperforms the much larger Gemini 3 Flash —

  • SWE-Bench Pro: 54.2% vs 49.6%
  • OSWorld-Verified: 74.0% vs 65.1%

3.6 Flash acts as the "master brain" to break down complex tasks, while Flash-Lite serves as a "distributed worker" to handle bulk execution.

One model handles planning, while a fleet of others handles execution, easily eliminating high-concurrency pressure.

In the official demo, the two models work in tandem: Flash-Lite generates 25 sets of highly available web design solutions in one go, so fast that it feels like "a batch of outputs appears as soon as you think of the task."

Flash-Lite is now available in the Gemini App, and will be gradually rolled out to Google Search.

Flash Cyber: Built for Vulnerability Patching, Not Accessible to Everyone

Gemini 3.5 Flash Cyber is a hardcore model dedicated exclusively to cybersecurity.

It has one single mission: identify vulnerabilities and remediate them.

The current awkward reality is that AI can find vulnerabilities faster than existing systems can patch them.

Google has integrated this model into CodeMender, its code security intelligent agent. By leveraging collaborative operation of multiple Cyber agents, it has set a new SOTA on the cybersecurity benchmark CyberGym, with costs far lower than larger, heavier models —

Using a more lightweight model to handle tasks that demand the highest precision.

However, this model is not available to the general public for the time being.

Gemini 4 Development Underway

The Most Aggressive Pre-Training in History

Even before Gemini 3.5 Pro is fully released, training for Gemini 4 has already begun.

These three consecutive launches are just the appetizer; the official announcement that stole the show is the real main event.

Google has confirmed that it has internally kicked off its most aggressive pre-training campaign in history, with the explicit target of developing Gemini 4.

Industry experts note that following Google's typical 6-month training cycle, we might see Gemini 4 as early as the end of this year.

For those of us who pay real money for tokens and build Agent systems, there is only one word that truly matters in today's announcement: reduction.

Prices are reduced, token consumption is reduced, and the cost of running an Agent is reduced.

As for the long-overdue flagship model and the newly started Gemini 4, they are still "futures" that are not yet available.

Models that you can actually use, at an affordable price, are the ones that deliver tangible value right now. No matter how powerful a flagship model is, it has to be deployed in the real world first.

References:

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/?utm_source=tw&utm_medium=social&utm_campaign=og

This article is from the WeChat Official Account "AIera", edited by Taozi, and republished by 36Kr with authorization.