The official version of DeepSeek V4 is here, new capabilities have been unveiled, and the king of cost-effectiveness has kicked off the battle.
In the afternoon of July 31, Phoenix Tech checked the official website of DeepSeek and found that the official version API of DeepSeek-V4-Flash has been launched for public beta. Different from the previous hierarchical logic of "Pro is strong while Flash is weak", this update releases a notable signal — the performance of the official Flash version in multiple Agent benchmark tests has approached or even exceeded the level of the V4-Pro preview version three months ago.
The official update log shows that the V4-Flash official version scored 82.7 on Terminal Bench 2.1, 54.2 on NL2Repo, 76.7 on Cybergym, and 70.3 on Toolathlon verified. While the V4-Pro preview version scored 67.9 on Terminal Bench 2.0.
It should be noted that Terminal Bench 2.0 and 2.1 are not test sets of the same version, so a direct comparison is not completely fair. However, a lightweight version with only 13 billion activated parameters achieving such scores in Agent capabilities already indicates that the optimization space in the post-training stage may have more leverage effect than simply stacking parameters.
In addition, DeepSeek also specifically stated that the current public beta is limited to APIs, and the latest capabilities are not available on the App and web pages for the time being. The official version of DeepSeek-V4-Pro will be released as soon as possible.
"Only post-training was re-conducted"
DeepSeek officially stated in the update log that "the model structure and size of DeepSeek-V4-Flash-0731 are consistent with DeepSeek-V4-Flash-preview, and only post-training was re-conducted."
According to DeepSeek's official technical report, V4-Flash has a total of 284 billion parameters and 13 billion activated parameters; V4-Pro has a total of 1.6 trillion parameters and 49 billion activated parameters. The two differ by an order of magnitude in model scale. If Flash can bring Agent capabilities close to the level of Pro through post-training, it means that for specific tasks, model scale is not the decisive factor — the weight of training methods and data quality is rising.
The official also specifically noted that the Code Agent task in this public benchmark test uses the DeepSeek Harness minimal mode as the framework for testing, with max gear, topp=0.95, temperature=1.0. This detail implies that DeepSeek may have made targeted optimizations at the Agent framework level, not just improvements to the model itself.
This is also the first time that DeepSeek's self-developed Harness has appeared with an official name. Previously, Liang Wenfeng once compared the implementation path of AGI to climbing stairs: the language model is the first step, last year the problem solved was CoT (Chain of Thought), this year's step is Agent, and the problem that must be solved after Agent is continuous learning — allowing the model to accumulate experience over time like humans, instead of being fed the complete context every time to work. Only after that comes the "singularity" of self-iteration and embodied intelligence.
"AI now has no lack of taste and intuition, what it lacks is the ability of continuous learning," he said, "Investors focus on Agent, while we focus on how to solve the learning problem."
Continuous learning sounds like a proposition at the model level, but its engineering foothold is precisely on Harness. DeepSeek's Agent Harness team was established in March this year, led by Cui Tianyi, a post-90s graduate of Zhejiang University majoring in Computer Science, who holds 6 ACM Asian Regional Gold Medals and worked at the top quantitative institution Jane Street for nine years. He joined DeepSeek in March this year, after which the team recruited a large number of talents in May.
It is revealed that DeepSeek's Harness plan will be launched simultaneously when the official V4 version goes online. As DeepSeek stated at the end of the log, "The official version of DeepSeek-V4-Pro will be released as soon as possible."
A Head-to-Head Confrontation at the Ecosystem Level
If the competition of large models in 2023 was "competing for general capabilities" and that in 2024 was "competing for long context", then the keyword for 2026 is undoubtedly Agent.
Agent capability — that is, the ability of models to plan autonomously, call tools, and perform complex tasks — is becoming a new yardstick to measure the strength of large models. From Terminal Bench (terminal operation), NL2Repo (code repository generation) to Cybergym (cybersecurity task) and SWE-bench (software engineering), a series of Agent benchmark tests are redefining what a good model is.
From a global perspective, the first echelon of Agent capabilities is still dominated by closed-source giants. According to data from third-party evaluation platforms such as benchlm.ai, the GPT-5.6 Sol and Claude Opus series rank at the top of most Agent benchmark tests. In the domestic camp, GLM, Qwen and other models are also catching up rapidly.
DeepSeek's substantial improvement of Flash's Agent capabilities this time has a clear strategic intention: to use cost-effective lightweight models to enter the huge market of Agent applications.
After all, Agent scenarios are far more sensitive to reasoning speed and cost than pure dialogue scenarios. For an Agent task that requires repeated tool calls and multi-round reasoning, if it runs on a flagship model, the cost may be several times or even dozens of times that of Flash. If Flash can reach a "sufficiently capable or even excellent" level in Agent capabilities, its cost performance advantage will be extremely competitive.
The two internal test sets announced by the official also have clear goals. DSBench-FullStack (internal full-stack development test set) scored 68.7, and DSBench-Hard (internal Coding Agent difficult problem test set) scored 59.6. This confirms DeepSeek's positioning from the side — to build Flash into the preferred base for developers and Agent applications.
A quiet ecological battle is also underway. The official V4-Flash natively supports the Responses API format and is specially adapted to Codex.
Responses API is a new generation API format strongly promoted by OpenAI. Compared with the traditional Chat Completions, it is more suitable for Agent scenarios — supporting more flexible tool calls, more complex multi-round interactions, and more fine-grained output control.
DeepSeek's native support for this format means that developers can migrate Agent applications developed based on the OpenAI ecosystem to DeepSeek at lower cost. This will be a head-to-head confrontation between the two at the ecosystem level.
The adaptation to Codex targets the vertical scenario of code generation and software development. Codex is OpenAI's model for the code field. DeepSeek's targeted adaptation is equivalent to directly launching a challenge in OpenAI's traditional advantageous field.
Putting these actions together, DeepSeek's ambition is not limited to being a "cheap substitute", but to establish its own ecological niche in the Agent era.
After 50 Billion Yuan of Financing: The Time Window Under High Valuation
This update came less than two months after DeepSeek completed its first round of external financing.
More than a month ago, DeepSeek completed its first round of external financing in nearly three years since its establishment, with a total amount of over 500 billion RMB, setting a single-round financing record in China's AI industry. The post-investment valuation is about 520 billion US dollars (about 3.5 trillion RMB). The investor lineup includes industrial giants such as Tencent, CATL, and JD.com, as well as multiple state-owned industrial funds.
In mid-July, DeepSeek intends to promote a new round of private placement financing, with a pre-investment valuation of about 710 billion US dollars (about 4.8 trillion RMB), an increase of about 37% compared with the 520 billion US dollar valuation after the first round of financing. The interval from 520 billion to 710 billion is less than six weeks.
Behind the high valuation is the market's recognition of DeepSeek's technical strength, as well as the bet on its commercialization prospects. But high valuation also means high expectations and high pressure.
According to reports from multiple media, Liang Wenfeng, the founder of DeepSeek, personally contributed about 200 billion RMB in the first round of financing, and firmly controlled the company through a special structure. The founder, who came from HF Quantitative, has insisted on investing with his own funds for nearly three years. Now his shift from "no financing and no listing" to actively embracing capital is a strong signal — DeepSeek is accelerating towards commercialization and IPO.
The launch of the official Flash version can be regarded as a technical delivery of DeepSeek under the support of capital. But the real test is yet to come: can the official Pro version be released as scheduled? Can the improvement of Agent capabilities be converted into real revenue? Under the siege of giants such as OpenAI and Anthropic, how will DeepSeek's cost-effective route give full play to its advantages.
The curtain of the Agent war has just been raised. With the substantial evolution of a lightweight model, DeepSeek has raised a new question for the industry — when the dividend of post-training is fully tapped, and when the improvement of efficiency begins to offset the gap in parameters, the competition logic of large models may be rewritten.
This article is from the WeChat official account "Phoenix Tech", author: Phoenix Tech, published with authorization from 36Kr.