Step aside, all new models: Google Gemini 4 is coming, and the largest large-scale model training project in the company's history has been launched.
IT Home News, July 22: In the early hours of today, Google released three new models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, with key improvements targeted at the token efficiency, response speed, and operational reliability required for Agent workflows.
▲Google launches Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber (Image source: X)
Among them, Gemini 3.6 Flash is designed for programming, knowledge work, and multimodal tasks. In the Intelligence Index from third-party evaluation firm Artificial Analysis, its task token consumption is 17% lower than that of Gemini 3.5 Flash, and its output token consumption in the DeepSWE test is reduced by up to approximately 65%.
Gemini 3.5 Flash-Lite is optimized for low latency and high throughput, with a generation speed of 350 output tokens per second, and it delivers significant improvements in Agent workflows compared to previous Flash-Lite generations.
Gemini 3.5 Flash Cyber is purpose-built to identify and remediate cybersecurity vulnerabilities. When paired with the CodeMender code security Agent, it reaches competitive performance levels with state-of-the-art models, offering capabilities on par with Mythos 5 and GPT-5.6 Sol, though it is not yet available to general developers.
Currently, on the Artificial Analysis leaderboard, Gemini 3.6 Flash achieves excellent speed test results, but its advantages in intelligence and cost are relatively less prominent: the model scores only 50 in Intelligence, trailing behind Claude Fable 5, GPT-5.6 Sol, Kimi K3, Grok 4.5, and GLM-5.2, and even Meta's Muse Spark 1.1 ranks above it. Its per-task cost reaches $0.50, making it more expensive than models from DeepSeek, MiniMax, GLM, Muse Spark, Grok, and others.
▲Comparison of Gemini 3.6 Flash against other models in capabilities, speed, and cost (Image source: Artificial Analysis)
As one of the original "top three" large model providers, Google's most prominent advantage in this release remains speed. Currently, Gemini 3.5 Flash-Lite and Gemini 3.6 Flash occupy the top two positions on the leaderboard for output speed.
▲Gemini 3.5 Flash-Lite and Gemini 3.6 Flash rank top two on the model output speed leaderboard (Image source: Artificial Analysis)
For users who have waited two months, this release feels somewhat like "waiting for the wrong product." Back in May, Google CEO Sundar Pichai announced at the I/O conference that Gemini 3.5 Pro would launch in June; now that July is already halfway through, users still have not received the 3.5 Pro model.
Regarding the long-delayed Gemini 3.5 Pro, Google states that the model is currently being tested with partners and will be released as soon as it is ready. At the same time, Google has given an early preview of Gemini 4, confirming that it has launched the largest pre-training task in the company's history. After keeping users waiting for two months, Google has put out two new "promises" that live up to expectations.
01. 3.6 Flash: Up to 65% reduction in output tokens, with marked improvements in programming and knowledge work
Gemini 3.6 Flash is refined based on feedback from developers and enterprise customers for the 3.5 Flash model. When handling multi-step workflows, the new model requires fewer reasoning steps and tool calls, while also reducing output redundancy.
The input pricing for 3.6 Flash is $1.5 per million tokens, and output pricing is $7.5 per million tokens. The output price is lower than that of 3.5 Flash, while the input price remains unchanged.
3.6 Flash drastically cuts the running cost of each Agent task. In the DeepSWE v1.1 test, 3.5 Flash consumes an average of about 276,000 output tokens per task, while 3.6 Flash reduces that number to approximately 97,000, a decrease of nearly 65%. In the Artificial Analysis Intelligence Index, the average output token consumption of the two models is about 28,000 and 23,000 respectively, representing a roughly 17% reduction in token usage for Gemini 3.6 Flash.
▲Task output token comparison between Gemini 3.6 Flash and Gemini 3.5 Flash (Image source: Google)
In the software engineering Agent evaluation DeepSWE v1.1, 3.6 Flash scores 49%, surpassing 3.5 Flash's 37% and 3.1 Pro's 12%. In the machine learning engineering test MLE-Bench, 3.6 Flash achieves a score of 63.9%, leading 3.5 Flash's 49.7% and 3.1 Pro's 42.6%.
In the knowledge work test GDPval-AA v2, the scores for 3.6 Flash, 3.5 Flash, and 3.1 Pro are 1421, 1349, and 965 respectively. In OSWorld-Verified, which tests Computer Use capabilities, the three models score 83.0%, 78.4%, and 76.2% respectively. Computer Use is now a built-in client tool within the Gemini API and Gemini Enterprise.
▲Other capability comparisons between 3.1 Pro, 3.5 Flash, and 3.6 Flash (Image source: Google)
Additionally, Google has showcased real-world application cases for the model: Gemini 3.6 Flash can analyze financial data and conference call transcripts more efficiently and accurately, execute code migration through multi-Agent collaboration, develop photo texture extraction tools for 3D workflows, and create interactive theme design tools in conjunction with the tldraw offline editor.
In a comparative code migration task test, Gemini 3.6 Flash takes 20 seconds for planning, less than 3.5 Flash's 24 seconds; the migration execution time is shortened from 2 minutes 8 seconds to 1 minute 20 seconds, cutting total task duration by approximately 34%. According to the case demonstration, 3.6 Flash generates more detailed, clearly structured code audit and verification documentation that covers targeted validation tasks for data models, API interfaces, business logic, UI components, and more.
▲Runtime comparison of Gemini 3.5 Flash and Gemini 3.6 Flash in code migration tasks (Image source: Google)
Several early customers have shared their feedback. Design software company Figma notes that 3.6 Flash strikes a strong balance between quality, speed, and cost, accelerating prototype exploration and iteration. Legal AI firm Harvey reports that the model averages a 12% speed improvement on document drafting and review tasks. Development tool company JetBrains states that its programming performance at low reasoning levels is 10% to 20% higher than the previous generation Flash model.
On the security front, 3.6 Flash enhances cutting-edge safeguards against chemical, biological, radiological, and nuclear risks, as well as misuse via cyberattacks. Google states that these measures improve the model's resistance to jailbreak attacks while minimizing false rejections for legitimate use cases.
02. 3.5 Flash-Lite: 350 tokens per second output, outperforming 3 Flash in select tests
Gemini 3.5 Flash-Lite is the fastest, lowest-priced model in the Gemini 3.5 family, targeted primarily at Agent search, document processing, and other tasks that prioritize low latency and high throughput.
According to data from Artificial Analysis, the model can generate 350 output tokens per second. Its input and output pricing are $0.3 per million tokens and $2.5 per million tokens respectively.
Developers can adjust the model's reasoning level based on their workload: select minimal or low levels for large-scale, low-cost tasks, and enable higher levels when handling multi-step sub-Agent tasks. 3.5 Flash-Lite also comes with the built-in Computer Use tool.
Compared to 3.1 Flash-Lite, 3.5 Flash-Lite sees its score in the Agent terminal programming test Terminal-Bench 2.1 rise from 31% to 54%; its score in the long-context test GDM-MRCR v2 increases from 60.1% to 72.2%; and its score in the knowledge work test GDPval-AA v2 goes up from 642 to 1140.
▲Comparison between Gemini 3.1 Flash-Lite and Gemini 3.5 Flash-Lite (Image source: Google)
3.5 Flash-Lite also outperforms the higher-tier 3 Flash in several tests. In SWE-Bench Pro, Gemini 3.5 Flash-Lite and Gemini 3 Flash score 54.2% and 49.6% respectively; in OSWorld-Verified, their scores are 74.0% and 65.1% respectively.
▲Comparison between Gemini 3 Flash and Gemini 3.5 Flash-Lite (Image source: Google)
US fintech company Ramp observes that Gemini 3.5 Flash-Lite strikes a strong balance between accuracy, latency, and cost in invoice extraction tasks.
▲US financial firm Ramp's evaluation of Gemini 3.5 Flash-Lite (Image source: Google)
03. Cybersecurity model: Designed for Agent integration, not yet open to general developers
The third model, Gemini 3.5 Flash Cyber, is fine-tuned from 3.5 Flash and purpose-built to identify, validate, and remediate cybersecurity vulnerabilities. Compared to larger, more resource-heavy models, it offers lower token pricing, making it suitable for large-scale code security scanning workflows.
The model will be paired with Google's code security Agent CodeMender. CodeMender runs multiple 3.5 Flash Cyber Agents in parallel, then aggregates the analysis results from different Agents into a single consolidated report.
In the CyberGym test, where CodeMender calls the model up to 5 times, 3.5 Flash Cyber scores 83.2%. Its performance on the test is close to leading-edge models from OpenAI and Anthropic: Mythos Preview in the Anthropic Agent scores 83.1%, GPT-5.6 Sol in the OpenAI Agent scores 83.6%, Mythos 5 in the Anthropic Agent scores 83.8%, and GPT-5.5-Cyber in the OpenAI Agent scores 85.6%.