The launch of its most powerful model has been delayed, with 200 billion US dollars wiped off its market capitalization overnight. Google has rolled out three "fuel-efficient" models in a row, sparing no effort to cut the cost of every single token to the extreme.
Last night, one day before Alphabet released its earnings report, Google launched three brand-new Gemini models in one go, all belonging to the "fast, low-cost" Flash series. The new releases focus on improving operational efficiency in complex agentic workflows and reducing token consumption, with the core feature clearly prioritizing efficiency over absolute performance.
In just one week, xAI's Grok 4.5, OpenAI's multiple versions of GPT-5.6, and Moonshot AI's Kimi K3 have been released one after another; meanwhile, Anthropic's Fable 5 has topped multiple leaderboards. In this "summer of being comprehensively overtaken," Google's current response is to offer cheaper models.
Trading Peak Performance for Efficiency and Lower Prices
According to Google, this new Flash series is designed to optimize the development of AI agents through stronger capabilities and higher cost-effectiveness. The series includes three models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber tailored for security scenarios.
Among them, Gemini 3.6 Flash is the "mainstay" in its flagship model lineup, focusing on achieving the optimal balance between quality and efficiency. Google states that compared to its predecessor, its knowledge cutoff date has been updated to March 2026, with improved capabilities in programming, knowledge-based tasks, and multimodal reasoning, while reducing average output token consumption by 17%.
At the same time, in some usage scenarios, Gemini 3.6 Flash can reduce output tokens by up to 65% compared to 3.5 Flash. In the post-"token budgeting" era, it indeed delivers notable performance in "token saving," which has become a highly meaningful metric. By reducing intermediate reasoning steps and tool calls, this model significantly compresses the overall operating cost of multi-step AI workflows. In terms of benchmark performance, Gemini 3.6 Flash shows obvious improvements in tasks such as code modification accuracy, machine learning research, computer operation, document parsing, and report generation. Enterprise customers including Heavia and Harvey have stated that the model's performance in multimodal tasks such as visual processing, chart analysis, and data extraction has been significantly enhanced.
Crucially, the model also has a lower cost, with an input price of $1.50 per million tokens and an output price reduced from $9 to $7.5 per million tokens. As people's awareness of AI usage costs continues to rise, this further enhances its market competitiveness. However, in terms of performance comparison, Gemini 3.6 Flash does not deliver a stunning surprise.
In most key benchmarks, it seems to lag behind Anthropic's Claude Sonnet 5 and OpenAI's GPT-5.6; even in tasks such as agentic coding, it has begun to be left behind by the latest Grok 4.5. Considering that Gemini 3.6 Flash is not much cheaper than its competitors, roughly on par with Grok 4.5 and GPT-5.6, it faces considerable difficulty in finding its own position in this model war.
Gemini 3.5 Flash-Lite is the "fastest and most cost-effective" model in the series, mainly targeting high-throughput agent search and large-scale document processing scenarios. This model can generate 350 tokens per second. Like 3.6 Flash, it further improves overall efficiency while maintaining the performance of its predecessor. Google states that Flash-Lite outperforms the old 3 Flash in multiple agent and programming benchmarks, including SWE-Bench Pro and OSWorld-Verified. Compared with its direct predecessor 3.1 Flash-Lite, its Terminal Workbench 2.1 score has increased from 31% to 54%.
The model also has a lower price, with $0.30 per million input tokens and $2.50 per million output tokens, providing a more cost-effective solution for high-load enterprise applications.
It can be seen that the two models Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are betting on the same strategic judgment: most AI usage scenarios do not require a "cutting-edge brain," but a model that is fast enough and cheap enough. Earlier this year, Google CEO Sundar Pichai said: "Enterprises are already burning through their annual token budgets, and it's only May." He said in an interview that the combined use of Flash series models can save enterprises more than $1 billion annually.
Currently, these two models are already available to Gemini Enterprise users and Gemini App users, and 3.5 Flash-Lite will also be integrated into Google Search in the near future.
Positioned as an "Affordable Alternative" to Mythos, With Strict Restrictions
The most targeted release is Gemini 3.5 Flash Cyber. Built on Gemini 3.5 Flash, this model is specially optimized for cybersecurity tasks to discover and fix software vulnerabilities. It is reported that Google has integrated it into CodeMender, the code security agent of Google DeepMind. This system runs multiple Flash Cyber sub-agents in parallel and aggregates the results into a unified report, representing Google's attempt to enter the dedicated cybersecurity model space.
Google positions Flash Cyber as a "cost-effective alternative" to large security models, claiming that it can reach the level of cutting-edge models in a key benchmark test at only a fraction of the cost. In the CyberGym benchmark, the model achieved a score of 83.2%, only about 2 percentage points lower than OpenAI's GPT-5.5-Cyber (85.6%), even though its model size is much smaller. It is worth noting that CyberGym is currently one of the few benchmarks that include comparative data of competing models, with very close scores across different models, and Google's advantage is mainly reflected in its lower token cost.
In another test, the Google Cloud Vulnerability Research Team used this model to scan public APIs and discovered a remote code execution vulnerability in just two hours, generating a valid exploit code that could bypass security defenses. In practical applications, Google's Big Sleep team used this model to detect critical vulnerabilities in Google Chrome and Safari, and its performance surpassed the standard Flash model as well as Anthropic's Claude Opus 4.6. When scanning code submissions for the V8 JavaScript engine, Flash Cyber found 55 confirmed independent issues, compared to 47 found by standard 3.5 Flash and 36 by Opus 4.6; 10 of these vulnerabilities were completely undetected by other models.
In addition, Anthropic's leading position in the field of AI security has also been challenged. Mythos is the "invisible rival" of this model, but its pricing is as high as $10 per million input tokens and $50 per million output tokens.
However, Google also acknowledges that such models are equally powerful in both "offensive" and "defensive" aspects, so access permissions must be strictly controlled. Currently, Flash Cyber is only available through CodeMender to government agencies and trusted partners in a pilot program for detecting and fixing security vulnerabilities. As for the CodeMender agent itself, a preview version (using standard Gemini models) has been made available via the Gemini Enterprise Agent platform, supporting multiple programming languages including C/C++, Go, Java, Python, Ruby, Rust, and TypeScript.
Its competitors have also adopted similar access restriction strategies for cybersecurity capabilities. Anthropic has implemented restricted access to its Claude Mythos, while OpenAI only opens some defensive capabilities of GPT-5.6 to verified users through its "Trusted Access" program.
Key Model Absent, Gemini 4 Reportedly Enters Pre-Training
Although Google has made progress on 3.6 Flash, what is more notable is actually "what was not released."
Currently, Google does not even have a single model in the top ten of public leaderboards. Gemini 3.5 Pro, the flagship model that Google originally planned to launch in June, is still in the testing phase. Wall Street is using this model as an important indicator to measure whether Google DeepMind under Google can maintain competitiveness with Anthropic and OpenAI. Last November, the launch of the Gemini 3 series models once showed that Google had returned to the forefront of the AI competition after years of falling behind, even prompting OpenAI to sound an internal "red alert" to accelerate the research and development of its GPT models.
On July 17, according to foreign media reports, the delay of Gemini 3.5 Pro is allegedly due to its failure to meet internal expectations, especially in programming capabilities, which has precisely become one of the most commercially valuable application scenarios for enterprise-level AI. In late June, the team tried to improve the model by updating the training data (the large-scale dataset that the model learns from), but the results were still disappointing. Affected by this news, Alphabet Inc.'s stock price fell by about 4.4% in a single day, with its market value evaporating by about $200 billion.
Ten current and former Google employees revealed that many people, including researchers and management, are worried that the company may lose its market advantage as competitors Anthropic and OpenAI launch models with capabilities surpassing Gemini. These anonymous insiders said that Google needs to coordinate stakeholders at multiple levels before releasing a model, while also striving to integrate AI into its huge product ecosystem, including Search, Maps, and YouTube. This complexity may also lead to a slowdown in the release schedule.
Google's response to this is to look further ahead. The company stated that it has launched "the most ambitious pre-training" it has ever undertaken, directly targeting Gemini 4. But this is more of an expression of intent than a realized capability. In other words, Gemini 4 is very likely already in the pre-training stage. Therefore, Gemini 3.6 is less of a major upgrade and more of a demonstration to remind the market that Gemini still exists "before the arrival of the truly major version."
The release of Google's three models this time first conveys the "efficiency first" strategy. Overall, this strategy is self-consistent: in the year when enterprises start to carefully calculate token costs, capture the mid-range market with "low cost + fast speed," while buying time for the flagship model. Moreover, this strategy is not only reflected at the software level. It is understood that Google is also developing self-developed chips to run Gemini at lower costs.
But whether this bet can succeed depends on two points: whether Gemini 3.5 Pro can be launched as soon as possible, and whether Gemini 4 will eventually be more than just a "promise in training."
References:
https://ai.google.dev/gemini-api/docs/latest-model
https://www.reuters.com/business/google-updates-lightweight-gemini-models-flagship-still-delayed-2026-07-21/
https://www.bloomberg.com/news/articles/2026-07-16/google-gemini-launch-delayed-as-tech-falls-short-of-internal-goals
This article is from the WeChat Official Account "AI Frontline", compiled by Hua Wei, and published with authorization from 36Kr.