Just now, the official version of DeepSeek-V4-Flash API has been launched for public beta testing.
On July 31, DeepSeek officially opened the public beta of the V4-Flash official version API to the public. The most notable highlight this time is — a substantial leap in Agent capabilities.
The full scores of 9 benchmark tests are released, with multiple indicators far exceeding the previously launched V4-Pro-Preview preview version, which has directly raised the ceiling of "coding agents" to a new level.
From Terminal Bench to DSBench-Hard, from cybersecurity confrontation to full-stack development, the official DeepSeek-V4-Flash version demonstrates remarkable potential as a "versatile Agent".
What does this mean? It means AI is no longer just "capable of writing code", but has started to work like a real engineer — opening terminals, analyzing repositories, calling tools, and completing end-to-end tasks.
9 Benchmark Test Scorecards: Data Speaks
This time, DeepSeek did not hold back, and directly presented the complete scores of 9 benchmark tests. Covering scenarios ranging from terminal operations to code repository analysis, cybersecurity to automated testing, its scope is extremely extensive.
Among them, Terminal Bench 2.1 led with a high score of 82.7, which means the model's task execution capability in the terminal environment is quite mature — it can understand complex command line operations, write scripts, and debug errors, delivering performance almost at the level of an "AI operations engineer".
Cybergym scored 76.7, demonstrating excellent cybersecurity confrontation capabilities. DSBench-FullStack (full-stack development test) and DSBench-Hard (Coding Agent challenge test) scored 68.7 and 59.6 respectively, proving that the model is not only capable of writing a single function, but also competent for the entire development chain from front end to back end.
Test Description: Transparent and Rigorous
DeepSeek also released detailed test conditions at the same time, and its transparency is highly commendable:
Note 1: For the Code Agent tasks in the public benchmark test sets, the official DeepSeek-V4-Flash uses the soon-to-be-released DeepSeek Harness minimal mode as the framework for testing, with the max gear, top_p=0.95, and temperature=1.0.
Note 2: DSBench-FullStack is an internally used full-stack development test set, and DSBench-Hard is an internally used Coding Agent challenge test set.
It is worth mentioning that the DeepSeek Harness minimal mode, as a test framework to be released independently soon, made its debut together with the test scores this time, which also implies DeepSeek's layout in building a standardized Agent evaluation ecosystem — not only training stronger models, but also providing more reliable evaluation tools.
Native Support for Responses API, Adapted to Codex
In addition to the leap in benchmark scores, the official V4-Flash also has important updates in engineering adaptation — it natively supports the Responses API format, and is specifically adapted to Codex.
This means developers can more naturally integrate V4-Flash into existing Codex workflows, reducing migration costs. The specific configuration method can be referred to the official DeepSeek documentation.
Model Structure Remains Unchanged, Post-Training Is the Key
A notable detail is that the model structure and size of DeepSeek-V4-Flash-0731 are completely consistent with the Preview version, and only post-training has been re-conducted.
In other words, the base model has not changed, what has changed is its "acquired training". This precisely shows that in the Agent capability track, there is still huge room for optimization of post-training strategies. With the same model architecture, more refined post-training tuning can bring a qualitative leap in benchmark tests — which has far-reaching implications for the entire industry.
Scope of This Upgrade: V4-Flash API Only
Upgrade Scope Description
Upgraded: DeepSeek-V4-Flash API interface
Not upgraded: DeepSeek-V4-Pro API
Not upgraded: App-side / Web-side models
To be released soon: DeepSeek-V4-Pro official version (will be released as soon as possible)
It should be noted that this upgrade only involves the API interface of DeepSeek-V4-Flash. If you are a user of V4-Pro, or a user who uses DeepSeek through the App/Web end, this update will not affect you for the time being.
As for the highly anticipated official V4-Pro version, DeepSeek said it "will be released as soon as possible". Considering that the Flash version has already shown such strong Agent capabilities, the performance of the official V4-Pro version is undoubtedly even more anticipated.
Industry Observation: Has the Agent Era Officially Arrived?
From Terminal Bench to Cybergym, from DeepSWE to Toolathlon, the names of these 9 benchmark tests themselves reveal a trend: AI capability evaluation is evolving from "writing a piece of code" to "completing an engineering task".
The official DeepSeek-V4-Flash version far outperforms the V4-Pro-Preview on multiple Agent benchmarks, sending a clear signal — the focus of competition in the post-training era has shifted from model scale to in-depth optimization of Agent capabilities.
When AI can skillfully operate terminals, analyze code repositories, conduct cybersecurity confrontations, and complete full-stack development tasks, we may be witnessing the arrival of a new era: AI is no longer just a coding assistant, but has become a real AI engineer.
And DeepSeek is running at the forefront of this path.
This article is from the WeChat Official Account "CSDN", organized by Meng Yidan, published with authorization from 36Kr.