HomeArticle

The only remaining advantage of Doubao in the United States is its speed...

硬核看板2026-09-07 20:28
Top-tier Flash model, but unfortunately it is still a Flash model.

Monthly version updates, I dare ask who else can do that?

Gemini 3.6 Flash in July, Gemini 3.7 Flash in August, Gemini 3.8 Flash in September. No one can match this update frequency...

The improvements of 3.6 and 3.7 Flash were negligible. I had completely lost confidence in Gemini before, and when I saw the news of 3.8 Flash release, my expectation was almost zero. However, I still tried 3.8 Flash, and I didn't expect that even as a Flash model, its capability improvement this time is quite remarkable.

01. Unit price unchanged, why does the task cost more?

The unit Token price of 3.8 Flash remains the same as before. For 3 consecutive generations of updates covering 3.6, 3.7 and 3.8, the price has not risen at all.

The price of our "US Soybean Pack" is lower than that of Doubao Seed 2.1 Pro, and even close to the peak price of DeepSeek V4 Pro.

That being said, the total cost still increased.

Because the biggest feature of 3.8 Flash this time is the increased number of reasoning rounds. The official stated that previous Flash models would rush to end the task before it is fully completed, and the new model solves this problem. The solution of Gemini 3.8 Flash is to let the model think a little longer, output a little more content, and run a few more rounds for the task.

This leads to a problem. Compared with previous generations of Gemini, although the unit Token price has not changed, under the same task, the cost still rises because the number of rounds required to complete the task generally increases. The output Token of 3.8 Flash in the high reasoning mode for a single task is 33% more than that of 3.7 Flash. Coupled with the increase of rounds for agent tasks, the average task cost has risen from $0.40 to $0.58.

Interestingly, Google's attitude towards the number of reasoning steps is not static at all.

When 3.6 Flash was released, the official emphasized avoiding detours: completing tasks with fewer reasoning steps and tool calls; when it came to 3.8 Flash, the official began to emphasize taking a few more steps to think.

But hey, is it really useful for us to let 3.8 Flash output more content like this?

With the same prompt, I asked Gemini 3.7 Flash and Gemini 3.8 Flash respectively to explain the sound production principle of electric bass to me. Both of them suggested me to pluck the strings for a test. 3.7 explained it very clearly: "Pluck the 4th string with force"; but 3.8 gave the description "Pluck the thickest 4th string violently with the fingers of your right hand"...

????

Hello? Did I spend all my money on this?

02. The "US Soybean Pack" can finally finish the task completely

If it can't even express itself clearly, where on earth is the progress of this method of running a few more rounds?

It turns out that it is designed for actual work. This iteration mainly brings improvement to Agent capability.

In the Arena Agent capability ranking, it rose from 32nd of the previous generation Gemini 3.7 Flash to 14th now.

After testing, the improvement effect is really quite obvious. Its capability has surpassed most lightweight models at the "Flash" level.

I asked 3.8 Flash to make an interactive simulation of a four-cylinder engine, which only took 1 minute, and all elements presented in the result are complete.

For comparison, 3.7 Flash directly failed to draw anything under the same prompt, and I specially added the note "need to perform verification" in the prompt. It seems that 3.7 really ignored this part completely just to finish the work as soon as possible...

After checking with GPT, I found that not only the initial dependency loading went wrong, but also the subsequent rotation process and ignition sequence were all written incorrectly. When I tested 3.7 Flash again, the familiar feeling that the "US Soybean Pack" writes code seriously but completely wrongly came back.

The same problem that exists in Gemini 3.7 Flash also appears in Flash models from other manufacturers.

Neither GPT-5.6 luna nor GLM-5.3 flash could load and render the engine piston normally, and the subsequent rotation and ignition processes were completely missing. What's more, the two models took 20 minutes and 35 minutes respectively, but the result still could not run at all.

However, DeepSeek V4 Flash ran successfully, but it could only run: the cylinder was not drawn clearly, and the effect of four-stroke ignition was not reflected at all.

03. Make up for insufficient quality with high speed

But if we talk about the biggest advantage of Gemini compared with other models on the market, it must be the speed.

The speed of the Gemini series is really fast now. For 3.8 Flash in high reasoning mode, the Token output per second reaches more than 300...

Among all reasoning models on the market, the top few in speed are all Gemini models.

This is related to Google's positioning for the Flash series models.

In the official product positioning, its main battlefield is Agent, and it is positioned as a "worker agent" that calls tools, runs continuously and completes tasks, rather than a "central agent" that thinks deeply, sets routes and makes overall scheduling. Therefore, it pursues high speed and low Token consumption, in fact, the performance of 3.8 Flash fully fits Google's positioning for the Flash series products.

I took my Pac-Man game to do another test.

I found that the speed of Gemini 3.8 Flash is really impressive. The whole game has complete functions and only took 29 seconds. That's way too fast...

I used the same prompt to test GPT-5.6 Sol xhigh, which also scored 59 points in the "Intelligence" dimension in Artificial Analysis. The result is much slower, it took 1 minute and 35 seconds, and the result is really not good, what is this frame rate trying to do...

,

Why can 3.8 Flash, as a lightweight model, still maintain good quality in many test tasks?

Because for it, high speed is the prerequisite for running more rounds.

I think Google's recent major updates to the Flash series are a very clever roundabout strategy.

In the past, when releasing large models, developers would first make a flagship model, and then distill a lightweight and fast Flash model from it. However, due to the continuous delay of Gemini 3.5 Pro, Google chose to go the opposite way: make up for insufficient quality with high speed. The 3.6 and 3.7 models were still focusing solely on improving speed, but for this 3.8 Flash, Google really made use of its high speed advantage to achieve better quality.

With this optimization method, the capability of Gemini 3.8 Flash can be regarded as a "flagship-level Flash" model, but after all, it is still not a Pro model. It is already September now, when will the Gemini Pro model originally scheduled for release in June be launched? I really can't wait any longer.

This article is from the WeChat Official Account "Hardcore Dashboard" (ID: yinghekb), written by Wang Rui, authorized