Liang Wenfeng refuses to defer to Elon Musk.
Overnight, Liang Wenfeng and Elon Musk had a head-on "collision".
On the night of August 12, DeepSeek V4-Pro was quietly launched in the API documentation first, before the official website banner and announcement "mysteriously" disappeared; it was not until 19:19 the next day that the official WeChat account officially announced its release.
From the debut of the preview version on April 24 to the launch of this official version, 111 days have passed. It was originally scheduled to be released in mid-July, and after several delays, it finally debuted in mid-August. Almost at the same moment, Elon Musk's SpaceXAI launched Grok 4.6.
The competition between two SOTA models on the same stage is not uncommon. Overall, DeepSeek V4-Pro has more advantages in specific fields such as AI safety (Cybergym) and workflow automation (AutomationBench); while Grok 4.6 performs better in tests such as the General Intelligence Index (Artificial Analysis) and software engineering (DeepSWE), and has significant advantages in cost efficiency.
On the morning of the 13th, when some developers questioned the capability of DeepSeek V4-Pro, DeepSeek dropped a blockbuster: Harness. Just the day before, Elon Musk also released the intelligent agent Grok Bot.
Intelligent agents have become the most intense part of this round of duel. Half an hour after DeepSeek Harness was launched, the number of GitHub stars exceeded 10,000, and exceeded 50,000 in 12 hours, building the Agent ecosystem with the concept of "everything is a plugin"; while Grok Bot takes the route of a cloud-based all-around assistant, focusing on 7x24 hours "follow-along learning".
But the duel is far from over. While the models are competing head-on, DeepSeek's self-built data center plan has been launched, trying to make up for the computing power shortcomings from the bottom layer. On the other side, Musk made a bold claim that within five years, SpaceXAI's computing power capacity will exceed the sum of the computing power of all other companies.
On one side is DeepSeek, which is focused on building its own computing power base and catching up by engineering efficiency, and on the other side is xAI, which holds a supercomputing cluster and has full ambition. The competition of foundation models tests hard power, the competition of intelligent agents fights for the next-generation application entrance, and computing power determines the upper limit of long-term development. The three major battlefields of models, intelligent agents and computing power are closely linked, and the three rounds of duel between Liang Wenfeng and Elon Musk have thus begun.
Foundation Models: Musk's Product Is Cheaper and Better
Claude Fable 5 launched by Anthropic is currently recognized as the ceiling of complex intelligent agent capabilities in the AI industry: with top capabilities such as long-link task processing and super-large engineering code reconstruction, it has long occupied the top position of the Agent track benchmark scores.
On the night when DeepSeek V4‑Pro and Grok 4.6 debuted simultaneously, the two models directly appeared under the nose of Fable 5 and launched impacts from different directions.
In terms of benchmark scores, the official published benchmark tests directly target the two currently recognized industry benchmarks, and DeepSeek V4‑Pro is generally on a par with Fable 5 in multiple indicators.
Grok 4.6 focuses on long-range intelligent agents, complex programming and knowledge work. The data released by SpaceXAI shows that it scored 1753 Elo in the real-world professional task evaluation GDPVal-AA v2, which is higher than the 1741 points of Claude Fable 5 Max. Its Artificial Analysis Intelligence Index reached 61.
After some users experienced it, they found that DeepSeek still maintains its usual advantage in speed, outperforming Grok 4.6 by a large margin. With the same prompt, it only took 16 minutes to complete the task, half the time of Grok, and the visual effect also exceeded expectations.
However, some users found after actual testing that DeepSeek V4-pro took longer time and consumed more tokens, but the performance was not very good. The most obvious problem is that near the line of sight, the water surface ripples should be larger, but DeepSeek hardly generated any ripples. In addition, the trajectory of the sailboat moving with the waves is also unnatural.
The evaluation results of the Artificial Analysis Intelligence Index show that V4-Pro scored 53 points, only 1 point higher than the Flash version, which made users dissatisfied.
In addition to performance, the biggest difference between the two is reflected in the pricing level.
The current pricing of DeepSeek V4‑Pro is $0.43 per million tokens for input, $0.87 per million tokens for output, and only $0.0036 for cache-hit input.
Taking the price of Fable5 as a comparison, the regular input of DeepSeek V4‑Pro is 23 times cheaper than Fable5, and the output is 57 times cheaper.
Then for Grok4.6, the input price is $2 per million tokens, which is 5 times cheaper than Fable 5; the output price is $6 per million tokens, which is more than 8 times cheaper than Fable 5.
When making a horizontal comparison between DeepSeek V4‑Pro and Grok4.6 directly, the price gap is also very wide. On the input side, the unit price of Grok4.6 is about 4.7 times that of DeepSeek V4‑Pro; on the output side, the price of Grok4.6 is nearly 7 times that of DeepSeek V4‑Pro.
However, after the official version was released yesterday, DeepSeek announced that with the official launch of the full V4 series of models, the API price will be updated and adjusted starting from August 17, adopting peak-valley pricing, and the price in off-peak hours will be half of the price in peak hours.
After the adjustment, compared with the peak hour price, the cache-hit input price of DeepSeek-V4 Pro has increased by 12 times, the cache-miss input price has increased by 3 times, and the output price has increased by 4.5 times. After the price increase, compared with Grok 4.6, the cost-effectiveness advantage of DeepSeek V4-pro during peak hours is no longer obvious.
However, even before the price increase, the cost advantage of DeepSeek V4 Pro was not so obvious. The development platform Composio conducted a comparative evaluation of 30 highly difficult intelligent agent tasks on the two models, and the results show that Grok is better than DeepSeek in terms of pass rate, speed and cost.
The Battle of Intelligent Agents: Completely Different Routes
While officially releasing DeepSeek V4-Pro, DeepSeek opened the DeepSeek Harness developer preview (v0.1) to global developers.
This is not unexpected. The announcement clearly states that the official version of V4-Pro has specially enhanced the Agent capability, and Harness is the first product that carries this capability.
DeepSeek Harness is a code intelligent agent team formed by DeepSeek in May this year. Its core goal is to develop a desktop intelligent agent programming product that competes with Anthropic Claude Code, and it is currently the most concerned team in DeepSeek.
Cui Tianyi, head of the Harness team, said on the "X" platform that the current 0.1 version is a preview version for Harness developers, which is still very imperfect, and he sincerely asks everyone to put forward valuable opinions.
Many internal beta users said that Harness has built-in multiple sub-agents and workflow capabilities, which can independently complete multi-step tasks such as requirement understanding, code generation, debugging and repair.
Its core function is to dynamically match different "toolkits" and "working modes" for the same large model to cope with various tasks. According to the complexity of the task, users can select different "Agent presets" for the model at the beginning of the session, just like equipping the same expert with different tools to complete tasks of different difficulties. Some developers commented that Harness is amazing.
Almost at the same time, SpaceXAI under Elon Musk also took action. On August 12, it launched the AI software "Grok Bot", claiming that it is a team of intelligent assistants that are online around the clock, which can accept and complete the actual work tasks assigned by users.
SpaceXAI said that Grok Bot was originally just an internal prototype, and it was unexpected that it would quickly become popular within the company. Teams have created Bots to handle sales outbound calls, marketing campaigns, daily office operations, vulnerability repairs and other work.
At present, Grok Bot is in the testing phase, and it is open to subscribers of SuperGrok Heavy, Cursor Ultra and Cursor Teams on platforms such as Windows and iOS starting from the release date.
DeepSeek Harness and Grok Bot have taken two completely different routes. The former is oriented to developers, with the core scenario of programming, while the latter is oriented to ordinary people who want to improve efficiency with AI. Actual tests found that reconstructing a web page with DeepSeek Harness only costs 3 yuan, but in long-range tasks, the response speed is relatively slow, requiring multiple rounds of interaction, and it is better at resource optimization for long text and multi-step tasks.
While Grok Bot has a fast response speed, its token cost is higher than Harness ($2 for input per million tokens, $6 for output per million tokens), and the total cost may be higher in long text scenarios.
Compared with DeepSeek Harness, SpaceXAI has far more cards than just Grok Bot. In June this year, SpaceX acquired the AI programming startup company Cursor for $60 billion.
In contrast, DeepSeek's pace is more restrained, focusing on independent research and catching up with engineering efficiency. It uses capital leverage to compress time on the one hand, and uses engineering iteration to gain development space on the other hand.
When both sides aim at intelligent agents at the same time, whoever can embed the model capability into the real work scenarios first will hold the entrance of the next generation of applications.
The Battle of Computing Power: Engineering Capability VS Cash Reserves
Computing power is the underlying fuel that determines who can go further, and it is also the ultimate competition in this duel.
Almost all model companies have encountered the situation of insufficient computing power. Kimi K3 under Moonshot AI refreshed the benchmark list at one stroke with its excellent performance, but because user requests far exceeded expectations, the computing power approached the bearing limit of the cluster, and the subscription was suspended 48 hours after its launch.
DeepSeek has frequently been on the hot search list due to service interruptions this year, and the behind reason is the serious mismatch between the surge in user scale and computing power expansion.
Liang Wenfeng once said that currently DeepSeek has about 20,000 H100 equivalent computing power cards, 12 to 18 months behind the United States, and the available number of domestic computing power cards is only 16,000, which is difficult to support the complete training iteration of super-large parameter models.
The limitation of computing power also makes DeepSeek more realistic in the choice of technical routes. Liang Wenfeng chose extreme engineering optimization, and the training cost of DeepSeek V3 is less than 5.6 million US dollars.
But this cannot meet the needs of users. DeepSeek has completed a Series A financing of 51 billion yuan, and a new round of financing is also in progress. Liang Wenfeng once said frankly at the investor meeting that the financing should be "converted into GPUs as soon as possible", and if the market permits, it is acceptable to purchase 20 billion yuan of computing power in 2026.
In contrast, Elon Musk, the richest man in the world, is much richer in computing power reserves. Grok4.5 was trained by tens of thousands of NVIDIA GB300 GPUs. The total computing power scale of its behind Colossus supercomputing cluster was about 200,000 H100 GPUs in 2024, which is 10 times that of DeepSeek's H100 equivalent computing power (about 20,000 cards).
The first 100,000 GPU supercomputing cluster of Colossus only took 122 days from start to launch, showing Musk's "money power" and execution ability in computing power construction. When DeepSeek is raising funds for a 20 billion yuan procurement plan, XAI has already put it into production.
Two weeks before the release of Grok 4.6, Elon Musk followed DeepSeek's official account on the "X" platform. Even his Grok admitted that this is equivalent to adding DeepSeek to the "list of opponents that must be closely watched". Before that, Musk even did not believe that DeepSeek R1 used so little computing power, and repeatedly said "it has at least 50,000 H100s".
This close-combat battle is still ongoing. Grok 4.6 is just an appetizer. Musk has clearly stated that Grok 4.7 will be launched in a few weeks, with the parameter scale expected to reach 2.1 trillion, and the goal is to "surpass all current models". And Liang Wenfeng has also been on the way to chase AGI (Artificial General Intelligence).
From models to intelligent agents, and then to computing power, the duel between Elon Musk and Liang Wenfeng has not yet been decided.
This article is from the WeChat Official Account "Tech Planet" (ID: tech618), authors: Ren Xueyun, Wang Lin, published with authorization from 36Kr.