Gemini 3.8 Flash Launches: Performance Approaches Opus 5, Base Price Remains Unchanged While Per-Task Costs Are Higher
Gemini 3.8 Flash arrived even faster than expected.
On August 13, Gemini 3.7 Flash was just launched; on September 2, Gemini 3.8 Flash has taken over. In the past 6 weeks, Google has updated the Flash model three times in a row.
The day before, Anthropic just released the "world's most advanced" Fable 5.1 and Mythos 5.1; OpenAI has also started warming up for its next-generation model Astra, which the official said will be launched "very soon". A new round of competition for cutting-edge models is kicking off. If Google doesn't come up with some "big moves" soon, there won't be many remaining model variants left for the Gemini 3 series.
Judging from Benchmark results alone, Gemini 3.8 Flash is already close to Opus 5 on some tasks.
In the long-horizon software engineering test DeepSWE v1.1, Gemini 3.8 Flash scored 73.7%, only 0.3 percentage points behind Opus 5's 74.0%; on Terminal-Bench 2.1, it even surpassed Opus 5's 89.1% with a score of 89.4%.
Google did not raise its price incidentally. The API input price of Gemini 3.8 Flash remains $0.75 per million Tokens, and the output price is $3.75 per million Tokens, which is exactly the same as the previous generation 3.7 Flash.
But this time, the low price is not as simple as it seems.
Google itself also reminds that 3.8 Flash will "work harder" than the previous generation — it will perform more reasoning, call more tools, and consume more Tokens. In the first round of tests by Artificial Analysis, the single-task cost of 3.8 Flash rose by about 40% compared to 3.7 Flash.
In other words, although the Token price has not increased, the cost of completing a task has become higher.
Google, it's enough to learn from top-tier models in terms of performance, don't learn this kind of thing from Anthropic...
Flash Encounters Flagship Models
According to the Benchmark results released by Google, the most obvious improvement of Gemini 3.8 Flash this time still focuses on Coding and Agent capabilities.
Compared with 3.7 Flash, its score on the long-horizon software engineering test DeepSWE v1.1 increased from 65.3% to 73.7%, a total jump of 8.4 percentage points, only 0.3 percentage points behind Opus 5's 74.0%.
The same situation applies to Terminal-Bench 2.1. In this practical command-line Agent test, 3.7 Flash previously scored 85.8%, while 3.8 Flash further increased to 89.4%, slightly exceeding Opus 5's 89.1%.
In addition, in Vals Finance Agent v2 for financial analysis tasks, 3.8 Flash scored 61.4%, higher than Opus 5's 58.6%; in Harvey's legal Agent test, 3.8 Flash got 10.0% while Opus 5 got 6.7%; on HLE-Verified, which measures multidisciplinary expert-level reasoning ability, the two scored 54.9% and 54.4% respectively.
Google therefore positions 3.8 Flash as the "smartest mainstay Flash model" at present, focusing on long-horizon software engineering, autonomous Agents and complex enterprise workflows.
At least in some Coding and Agent tasks, Gemini 3.8 Flash has already caught up with Opus 5.
However, it has not yet achieved full overtaking for the time being.
On Terminal-Bench 4.0, which is more difficult and emphasizes general Agent capabilities, 3.8 Flash only scores 19.1%, while Opus 5 reaches 51.8%; in the GDPVal-AA v2 knowledge work test, 3.8 Flash gets 1545 Elo while Opus 5 reaches 1824; on OSWorld 2.0 that tests computer operation capabilities, the two score 59.0% and 75.4% respectively.
The results given by independent evaluations are also closer to this judgment: Artificial Analysis ran the three reasoning intensity levels of Gemini 3.8 Flash, namely low, medium and high. The high level scored 59 points on the Artificial Analysis Intelligence Index, while Opus 5 max currently scores 63 points.
By the way, Fable 5.1 max, which was released yesterday, has reached 66 points, and the top seven positions on the current list are all occupied by different reasoning intensity levels of Anthropic models.
Therefore, to be precise, Gemini 3.8 Flash can only be regarded as a Flash model that has begun to compete across levels on some flagship tasks.
In terms of speed, it maintains the advantages of "Flash". In tests conducted by Artificial Analysis through Google API, the output speed of 3.8 Flash high reaches 304.6 Tokens per second, ranking first among the 195 models in its comparison group.
In terms of core product specifications, 3.8 Flash is exactly the same as the previous generation. It still has a context window of about 1 million Tokens and a maximum output of 65,536 Tokens, and the main supported tool capabilities remain unchanged. Google's model card directly states that 3.8 Flash is built on 3.7 Flash, which is a continuous iteration rather than a replacement of the underlying model.
Along with Gemini 3.8 Flash, Gemini 3.8 Flash Cyber has also been upgraded.
The new 3.8 Flash Cyber surpasses the previous generation 3.5 Flash Cyber in the CyberGym vulnerability discovery test; in Google's internal real vulnerability test covering 20 programming languages, its success rate exceeds 70%; it reaches 47.2% in the CWE-Bench automatic vulnerability patching test, surpassing GPT 5.6 Sol.
Google and Anthropic have different practices in terms of "security models". Fable 5.1 and Mythos 5.1 are essentially the same underlying model, and the scope of access is only distinguished through different cybersecurity and biosafety restrictions; Google has created a separate Cyber branch since 3.5 Flash, which is specially fine-tuned for vulnerability discovery, verification and patching on the basis of Flash, while adopting more relaxed cybersecurity restrictions for the Cyber version.
Interestingly, this kind of special training has begun to affect the main model in turn.
Google claims that 3.8 Flash and 3.8 Flash Cyber share the same underlying intelligent core, and part of the improvement in Coding and Reasoning of this generation comes from training in the highly difficult field of cybersecurity.
Model Pricing is Shifting from Token Price to Task Cost
If you only look at the API price list, Gemini 3.8 Flash is almost a free performance upgrade.
Its input price is still $0.75 per million Tokens, output price is $3.75 per million Tokens, and cache reading price is $0.075 per million Tokens, which is exactly the same as 3.7 Flash. The performance has improved, and the unit price of Tokens has not increased at all.
But when you actually run it, the bill is not calculated this way.
In Artificial Analysis's tests, the average cost for Gemini 3.8 Flash high to complete an Intelligence Index task reaches $0.58, while the previous generation 3.7 Flash only costs about $0.40.
In other words, with the same Token price, the single-task cost has increased by 40%.
The reason is simple: 3.8 Flash has become better at "thinking" and also consumes more Tokens.
Artificial Analysis found that 3.8 Flash high generates an average of about 48,000 output Tokens per task, an increase of 30% compared to 3.7 Flash; in Agent tests, it will also perform more interactions. In the end, although Tokens have not become more expensive, more Tokens need to be purchased to complete the same set of tests.
It is worth noting that this is not a problem encountered by Google alone. Anthropic has just demonstrated a more extreme case with Fable 5.1.
Fable 5.1 did not adjust the regular input and output prices this time, which remain $10 and $50 per million Tokens, but Anthropic directly reduced the cache reading price from $1 to $0.25, a drop of 75%.
For Agents, this is a very substantial price cut. Because Agents need to continuously read codes, documents and historical conversations that have already entered the context during long-time work, cache reading often accounts for a large amount of input Tokens. Anthropic estimates that this adjustment alone can reduce the cost of typical workloads by about 25%, and the cost of highly Agentized tasks can be reduced by up to about 45%.
However, when Artificial Analysis ran the test, the situation was reversed again.
Although Fable 5.1 max topped the Artificial Analysis Intelligence Index with 66 points, it costs an average of $3.76 to complete an Intelligence Index task, while the previous generation Fable 5 max only costs $3.14.
The cache reading price has been reduced by 75%, but the single-task cost of Fable 5.1 is 20% more expensive instead.
The problem still lies in Token consumption, but Fable 5.1 is not "too thoughtful", it is "extremely verbose".
Artificial Analysis found that the output Tokens generated by Fable 5.1 max are about 1.7 times that of Fable 5. The cache price reduction has saved about $1.40 per task. Without this 75% cache discount, its single-task cost would even reach about $5.16.
Thus there is a very counterintuitive result: