HomeArticle

Liang Wenfeng goes head-to-head with Elon Musk, and DeepSeek V4 Pro is putting pressure on the world's most powerful AI model.

凤凰网科技2026-08-13 09:51
We want both high performance and excellent cost-effectiveness.

In the past, high-performance models were far ahead, and even models with poorer performance could capture a certain market share relying on low prices. However, after the release of cost-effective high-performance models such as the DeepSeekV4 series and Gork4.6, this simple development path is likely to officially come to an end.

After much anticipation, DeepSeek's most powerful model has finally been unveiled.

Late at night on August 12, the API documentation on DeepSeek's official website was quietly updated — the model version was switched from the V4-Pro preview version to "DeepSeek-V4-Pro-0813".

Developers were still the first to notice the change. Shortly after, an evaluation comparison table leaked from the official DeepSeek group, lighting up the entire late-night AI circle.

It should be noted that the highly anticipated Harness still remains a mystery. Phoenix Tech's search of platform information found that the official WeChat account "DeepSeek Harness Team" was officially registered on July 6 this year, with the certification subject being "Beijing Deepseek Artificial Intelligence Basic Technology Research Co., Ltd.".

On August 1, Cui Tianyi, head of the internal Harness team at DeepSeek, publicly solicited Harness beta testers worldwide, prioritizing developers with experience in the open-source Agent Harness project. Applicants were required to submit their code hosting platform accounts and representative works.

According to developers who have tested it, "After participating in the internal test and experiencing it in depth for 2 days, I would call DeepSeek Harness the most powerful in the universe".

Just six days before the official release of V4 Pro, DeepSeek announced that it plans to raise the overall pricing of its API services in the near future, with "an expected significant increase". And the extreme cost-effectiveness of DeepSeek V4 Flash has just left a deep impression overseas.

On one hand, its performance is catching up with the world's top models, and on the other hand, it has announced a significant price hike. This company that has swept the market with cost performance is telling a completely different new story.

01

A 0.1-point gap, competing head-to-head with the world's strongest models

This widely shared evaluation table shows that on Terminal Bench 2.1, a test that measures the ability of AI agents to complete complex tasks in real terminal environments, the official V4 Pro scored 87.9 points. The world's top Anthropic Fable 5 scored 88.0 points, leaving a gap of only 0.1.

As a reference, the score of the previous V4 Pro preview version was 72.1 points, marking a 15.8-point jump in three months.

In addition, it surpassed its competitors in two more benchmarks. On Cybergym, the AI security agent benchmark, V4 Pro scored 83.3 points, narrowly edging out Fable 5's 83.1 by 0.2 points. In the AutomationBench workflow agent test, it also won with a score of 31.8 against 29.1. On DeepSWE, which measures software engineering capabilities, its score jumped from 12.8 points in the preview version to 62.7 points, nearly five times the original figure, exceeding the 58.0 points of Anthropic's previous flagship Opus 4.8.

It should be noted that these data are currently circulating in DeepSeek's official groups and communities, and the official has not yet released the official update log for V4-Pro-0813, so strictly speaking, they should be regarded as unofficial disclosures. However, cross-verification of experiences from multiple sources confirms that the performance improvement is real and tangible.

Agents are the key word to understand this upgrade. Different from traditional large models that compete in knowledge Q&A and mathematical reasoning, tests such as Terminal Bench and DeepSWE examine whether AI can actually get work done — call tools, perform multi-step tasks, write and modify code, and keep advancing in long chains until the results are delivered. This is exactly the core threshold for large models to evolve from chat toys to production tools.

According to DeepSeek's official API documentation, the official V4 Pro supports 1M (million) Token context length, maximum output of 384K Token, supports both thinking and non-thinking modes, and is compatible with both OpenAI and Anthropic API formats. This means developers hardly need to modify their code to switch applications that originally called Claude to DeepSeek.

The timing of this update is also very interesting. Just a few hours before the launch of V4 Pro, SpaceXAI under Elon Musk released Grok 4.6, which topped the GDPVal-AA v2 assessment for real knowledge work with 1753 Elo, surpassing Fable 5's 1741 and GPT-5.6 Sol Max's 1728. Two cost-effective models were released on the same night, both targeting the same goal: enabling AI to stably deliver usable results in long tasks. The showdown between Liang Wenfeng and Elon Musk, both launching products with high performance and high cost performance, means that OpenAI and Anthropic are probably the ones who truly feel the pressure.

At present, Fable 5 is priced at 50 USD per million output Tokens, GPT-5.6 Sol Max at 30 USD, Grok 4.6 at 6 USD, while DeepSeek V4 Pro is only 0.87 USD — about one fifty-seventh of Fable 5's price.

Behind the 0.1-point score gap lies a 57-fold price difference. This equation is enough to make every developer re-examine their technical selection.

02

The confidence behind a price hike notice

The performance of DeepSeekV4 Pro this time also partly explains the confidence behind the price hike announced six days ago.

On August 6, many developers saw a prompt when logging into the DeepSeek API backend: "We plan to raise the overall pricing of DeepSeek API services in the near future, with an expected significant increase. Please arrange your usage reasonably. The specific plan will be subject to the official notice."

This is also DeepSeek's first announcement of an overall API price hike, rather than the previous partial adjustments for peak hours.

The reason why the market was shocked is that DeepSeek's image over the past year has been exactly that of a "price butcher".

Sorting out its pricing timeline this year: when V4 was released in April, the V4 Pro API was launched with a 75% discount; after the discount ended on May 31, the price was directly and permanently reduced to a quarter of the original price, with 3 yuan per million Tokens for input (cache miss) and 6 yuan per million Tokens for output, a 75% drop from the launch price; on June 29, peak-valley billing was introduced, doubling the price during peak hours on workdays.

Only more than two months separated the permanent price cut and the announcement of a major price increase.

DeepSeek certainly has its own difficulties. Data from OpenRouter shows that in the week from July 27 to August 2, DeepSeek V4 Flash ranked first in the world with 7.11 trillion Token calls; according to statistics from the open-source project OpenCode, on August 1 alone, V4 Flash processed 8 trillion Tokens. On August 4, the model even suffered from insufficient capacity due to unprecedented access volume.

Call volume is both honey and a burden. Behind every API call is real computing power consumption. Long-term low-price high-volume operations are continuously eating into cash flow.

However, if the price hike is only interpreted as "unable to bear the cost", people may underestimate DeepSeek's intention.

A detail that is easy to overlook is that just two months ago, on June 16, DeepSeek completed its first round of external financing since its establishment, raising a total of more than 50 billion RMB (about 7.4 billion USD), with a post-investment valuation exceeding 50 billion USD, setting a single-round financing record for China's AI industry. Among them, founder Liang Wenfeng personally contributed about 20 billion RMB, making him the largest investor. The latest news shows that DeepSeek is already pushing for the second round of financing with a guaranteed minimum of 10 billion RMB.

Holding a large amount of capital but choosing to raise prices is more like an active pricing strategy shift, rather than a passive cost pass-through.

The research report released by Morgan Stanley on August 9 puts forward three driving factors: first, the strong demand for the V4 model supports stronger pricing power; second, manufacturers need to strike a balance between market share and gross margin to sustain investment in cutting-edge model R&D; third, the wider use of domestic chips may lead to higher reasoning costs.

The report believes that model intelligence, rather than price, is the ultimate barrier for large model competition. When V4 Pro is only 0.1 points behind Fable 5 on Terminal Bench, DeepSeek has been qualified to charge a premium for this "sufficiently close" performance.

03

The endgame signal of the price war

DeepSeek's price hike is not an isolated case.

Looking at the entire industry, this price hike trend has already emerged. According to Morgan Stanley's statistics, taking ByteDance, Alibaba, Baidu, Tencent, Zhipu, Moonshot AI, MiniMax and DeepSeek as samples, the average API output price of Chinese large models has rebounded from about 12.2 yuan per million Tokens in Q1 2025 to 21.9 yuan per million Tokens in Q2 2026.

The inflection point has already appeared. In the first quarter of this year, Zhipu's API pricing was raised by a total of about 83% compared to the end of last year, but the call volume increased by 400% instead — the simultaneous rise in volume and price shows that developers are willing to pay for more powerful models. In July, after Moonshot AI released Kimi K3, the number of requests within 48 hours exceeded expectations, approaching the cluster's carrying limit, so it directly suspended new C-end user subscriptions to fully expand capacity. Tencent raised its model API prices twice between March and April, with some increases as high as 463%. MiniMax, Alibaba Cloud and other manufacturers have also successively tightened free quotas and adjusted billing methods.

According to 36Kr's report, the domestic large model track has been mired in Token price wars for the past two years, with manufacturers seizing developers and enterprise customers by lowering interface quotations. However, in the second half of 2026, the supply and demand of computing power, chip procurement costs and operating costs continue to rise, and the model of relying solely on low prices to scale up is unsustainable.

The logic is actually not complicated. Some industry analysis points out that the computing power consumption of Agent tasks is 10 to 100 times that of ordinary conversations, and when models truly enter the production processes of enterprises — writing code, doing analysis, running workflows — customers' sensitivity to prices will give way to requirements for reliability and capability. Zhipu's 83% price hike still cannot meet demand, which has already illustrated this point.

Morgan Stanley summarizes this as a flywheel: stronger models bring more revenue, more revenue supports greater investment, and greater investment further improves model capabilities. Leading manufacturers can launch low-cost lightweight versions at any time to cover the mass market, but it is completely a different level of difficulty for mid-tier manufacturers to break through to the top level.

This may be the real footnote to DeepSeek's late-night version update. It first used V4 Pro to prove that it "can get work done", getting the ticket to compete on the same stage with the world's top models; then it used a price hike notice to test the market's willingness to pay for "Chinese intelligence".

Of course, new challenges have also come. After the price hike, will small and medium-sized developers who are extremely cost-sensitive migrate? How large exactly is the "significant increase"? In the context where players at the same tier such as Grok 4.6 are still attacking fiercely with low prices, how long can DeepSeek's pricing power last?

These questions can only be answered after the price hike is officially implemented. But one thing is certain: when the player best at fighting price wars starts to lay down their arms, the competition logic of China's large model industry has turned a new page.

This article is from the WeChat official account "Phoenix Tech", author: Dale, editor: Dong Yuqing, published with authorization from 36Kr.