Just now, Gemini 3.6 Flash was officially released, but netizens are laughing even louder.
If you've been using Gemini since last year, it feels like watching your own brother slowly develop Alzheimer's disease.
This is a recent playful jab from netizens about Gemini.
Logically, when a company releases a new model, it usually does so to debunk such mockery. But right after Google launched three new models in one go, netizens not only didn't take back that comment, they laughed even harder. It's almost like the whole world had no faith in you, and you just had to go out of your way to prove everyone right.
The three models are 3.6 Flash, 3.5 Flash-Lite, and a security-focused 3.5 Flash Cyber. The official blog's wording is eerily familiar: "more efficient", "smarter", "built for large-scale AI agents" — every buzzword is there.
However, the real-world experience is a total contrast.
The Star of the Show: What Exactly Makes 3.6 Flash Stand Out
Let's start with the flagship, 3.6 Flash. The official positioning is just two words — "workhorse", a model built to get things done.
3.5 Flash was announced at the I/O conference back in May this year, so this is essentially a minor version iteration.
Its biggest selling point is token efficiency.
According to data from the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash, and in scenarios like DeepSWE (a benchmark from Datacurve), the savings can reach up to 65%. Google also claims it requires fewer inference steps and tool calls when running multi-step tasks. In other words, for the same workload, there's less unnecessary verbosity, fewer detours, and a lighter bill.
Prices have indeed been reduced: $1.5 per million input tokens, and $7.5 per million output tokens — note that the previous generation 3.5 Flash charged $9 for outputs, so this is a direct cut to $7.5. Faster and more cost-effective, the cost of individual agent tasks has been driven down significantly.
On the benchmark testing front, Google has laid out a full array of data.
Coding capabilities are the key focus. DeepSWE performance has risen from 37% to 49%, with the company stating there are fewer unnecessary code modifications, shorter execution loops, and generated code that better aligns with production environment requirements. The improvement on MLE Bench for machine learning research is even more notable, jumping from 49.7% to 63.9%.
Computer use capabilities have also improved, with OSWorld-Verified scores climbing from 78.4% to 83%. Moreover, "computer use" is now a built-in tool in the Gemini API and enterprise edition, designed for out-of-the-box usability.
For knowledge work tasks, GDPval-AA v2 scores have increased from 1349 to 1421. Clients including Hebbia and Harvey report that it delivers exceptional performance on multimodal tasks like document parsing, chart data analysis, and report drafting. Figma and JetBrains have also publicly endorsed its capabilities.
Oh, and there's a small but meaningful update: the knowledge cutoff date has finally been pushed from January 2025 to March 2026. That outdated reference point is now a thing of the past.
On the security front, Google says 3.6 Flash features an upgraded Frontier Safety protection system, which focuses on preventing misuse in two key areas: CBRN (chemical, biological, radiological, nuclear) threats and cyberattacks. It offers stronger jailbreak resistance while minimizing false positives and unnecessary rejections of legitimate requests.
· Punching Above Its Weight: The Affordable, High-Value 3.5 Flash-Lite Now Supports Adjustable Inference Modes
The second model is 3.5 Flash-Lite, positioned one tier below 3.6 Flash. It prioritizes "speed" and "low cost", specifically optimized for high-throughput, low-latency tasks such as agent search and document processing.
It's the fastest model in the 3.5 series, clocking 350 tokens per second according to Artificial Analysis measurements. The price is extremely competitive — $0.3 per million input tokens and $2.5 per million output tokens. Google claims its quality is a significant step up from the 3.1 Flash-Lite released back in March.
One very flexible design feature: it supports adjustable "inference gears". For simple tasks, you can run it on the lowest gear for maximum speed and savings; for multi-step subtasks that require deeper reasoning, you can crank up the thinking level. Computer use functionality is also built in as a native tool.
Compared to its predecessor, the improvements are quite substantial: Terminal-Bench 2.1 scores for coding and agent tasks have risen from 31% to 54%; GDM-MRCR v2 for long-context tasks has gone from 60.1% to 72.2%; and GDPval-AA v2 for real-world task execution has skyrocketed from 642 to 1140.
The most interesting part is that this "younger sibling" model has actually outperformed the older 3 Flash in many agent and coding tasks.
For example, it beats 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%). Punching far above its weight class, it's genuinely impressive — meaning that for workloads that previously ran on 3 Flash, there's now a faster, more powerful alternative.
Google has shared several use cases: extracting product features from massive e-commerce datasets, acting as a fast assistant to 3.6 Flash to generate 25 web design drafts in one go, batch processing translated receipt summaries, and iteratively building small games through trial and error. Clients including Ashler, Palo Alto Networks, and Ramp have also praised its perfect combination of "speed, intelligence, and cost efficiency".
The final model, 3.5 Flash Cyber, takes a completely different approach — it's purpose-built to identify and fix security vulnerabilities in code.
Google's logic here is quite interesting: AI can now find vulnerabilities faster than existing systems can patch them. Since vulnerabilities are piling up faster than they can be addressed, the company decided to use the affordable, efficient Flash model to batch-process these fixes.
This model is fine-tuned on top of 3.5 Flash, and works in tandem with Google's CodeMender tool. CodeMender runs multiple Flash Cyber agents that collaborate to analyze issues and compile a final consolidated report.
On the well-known industry CyberGym benchmark, Google claims it has reached top-tier performance levels, while delivering better token efficiency than much larger models.
That said, most of us won't be getting access to this model anytime soon.
Given that vulnerability-hunting capabilities are a double-edged sword, even the traditionally straightforward Google has followed Anthropic's lead in emphasizing security guardrails. The model is heavily restricted — it's only available in limited quantities to "trusted partners" through a closed beta program. The goal is to give defensive security teams a head start to patch vulnerabilities before they can be exploited by bad actors, while preventing malicious use of the technology.
Where's the Promised 3.5 Pro? Nowhere to Be Found
At this point you might be asking: after releasing all these Flash models, where is the actual flagship 3.5 Pro?
The official line is that it's "being tested with partners and will launch when ready". But the story behind this statement might not be so pretty.
According to reports from Bloomberg and other outlets, the code generation performance of 3.5 Pro has consistently fallen short of internal expectations. Google specifically updated its training data in late June to address this coding weakness, but the results still didn't improve. There are even rumors that Google completely scrapped its original training plan and restarted the entire process from scratch.
Meanwhile, Logan Kilpatrick, Google AI's chief evangelist, has completely skipped over 3.5 Pro in public communications, shifting the focus of announcements to Gemini 4, claiming that they have launched the most ambitious Gemini 4 pre-training project in history, and stating that the progress so far is "exciting".
While this sounds uplifting, when viewed against the backdrop of a delayed flagship model, it comes across as little more than a distraction and a way to buy time with empty promises.
Beyond these behind-the-scenes stories, even looking at the newly released Flash models themselves, netizens are far from universally impressed.
Independent evaluation firm Artificial Analysis stated that while the two new models do cut single-task completion times in half and improve token efficiency, with Flash-Lite seeing an 11-point jump in its intelligence index — the intelligence level of 3.6 Flash is basically unchanged from 3.5 Flash, treading water.
The price isn't low enough to justify the negligible capability gap, and the performance isn't high enough to explain its costs.
One X user complained: 3.6 Flash scores exactly the same as 3.5 Flash on Artificial Analysis, and lags behind a whole host of competitors including Meta Spark 1.1, GLM-5.2, 5.6 Luna, Sonnet 5, Grok 4.5, and 5.6 Terra. They didn't mince words, calling it straight-up terrible (sarcasm implied).