Liang Wenfeng launches a surprise offensive against Elon Musk, pitting DeepSeek V4 Pro against Grok 4.6, and the first test delivers absolutely explosive performance.
Today, this day is absolutely surreal!
On one side, Liang Wenfeng finally released the official version of DeepSeek V4 Pro in the wee hours of the morning.
On the other side, Musk launched his next-generation flagship Grok 4.6, which is positioned to deliver low cost and superior performance.
The two tech giants released their new products at the exact same time, instantly ramping up the sense of rivalry to the maximum.
They are targeting almost the exact same thing —
Enabling the model to continuously call tools, modify code, and verify results in long-running tasks, and finally deliver a fully functional, usable end product.
DeepSeek V4 Pro is here
Musk's Grok 4.6 enters the arena
The most remarkable achievement of the official DeepSeek V4 Pro is that it took the absolute first place in two benchmarks, directly overtaking Fable 5.
It scored 83.3 points in the cybersecurity agent test CyberGym, surpassing Fable 5's 83.1 points and Opus 4.8's 78.3 points.
It scored 31.8 points in the automation task benchmark AutomationBench, outperforming both Fable 5's 29.1 points and Opus 4.8's 27.2 points.
The most head-to-head competition took place on Terminal-Bench 2.1.
DeepSeek scored 87.9 points, exceeding Opus 4.8's 85.0 points, and was only 0.1 point short of Fable 5's 88.0 points.
In several other high-difficulty Agent tests, DeepSeek achieved a clear lead over Opus 4.8.
It scored 60.0 points in the "Final Human Exam" with tool access, surpassing Opus 4.8's 57.9 points and continuing to close in on Fable 5's 63.0 points.
Even in the software engineering agent field, which used to be its weakest segment, the performance of DeepSWE skyrocketed from 12.8 points in the preview version to 62.7 points, nearly 4.9 times the original score.
This result not only exceeds Opus 4.8's 58.0 points, but also leaves only 7.3 points between it and Fable 5's 70.0 points.
At the same time, Musk also put Fable 5 and GPT-5.6 Sol under huge competitive pressure.
Grok 4.6 first caught up with GPT-5.6 in terms of comprehensive capabilities, and then continuously overtook it in coding tests.
It scored 61 points on the General Intelligence Index, only 1 point behind Fable 5's 62 points.
It reached 69.9% on CursorBench, surpassing GPT-5.6's 67.2%, and only 0.6 percentage points behind Fable 5's 70.5%.
It also reached 61.3% on FrontierCode, outperforming GPT-5.6's 60.6% and continuing to chase Fable 5's 63.6%.
When it comes to knowledge work scenarios that are closer to real workplace delivery, Grok 4.6 achieved a full reversal, taking first place in all three benchmarks.
It scored 1753 Elo in GDPval-AA v2, exceeding both Fable 5's 1741 and GPT-5.6's 1728.
It scored another 1577 Elo on AA-Briefcase, outperforming Fable 5's 1574 and GPT-5.6's 1502.
It reached 15.8% in the professional legal task benchmark Harvey LAB, while Fable 5 only scored 11.3% and GPT-5.6 scored a mere 2.5%.
What is more striking is that DeepSeek V4 Pro and Grok 4.6 are not only approaching the top-tier models of OpenAI and Anthropic in terms of performance, but also jointly driving down the price of cutting-edge AI intelligence.
For every million output tokens, Grok 4.6 costs 6 USD, GPT-5.6 Sol costs 30 USD, Claude Opus 5 costs 25 USD, and Fable 5 costs 50 USD.
While DeepSeek only costs 0.87 USD —
That is roughly 1/7 of Grok 4.6's cost, 1/35 of GPT-5.6 Sol's cost, 1/29 of Claude Opus 5's cost, and 1/57 of Fable 5's cost!
First public real-world test
DeepSeek battles Grok head to head
Now, the first batch of real-world test results has been released. DeepSeek V4 Pro and Grok 4.6 have finally moved from the benchmark score leaderboards to real task scenarios.
In the first round, we directly pushed DeepSeek to its limits.
Using only one single prompt, it built a complete 3D interactive Earth from scratch in the browser.
All basic interactions including dragging, rotating, and zooming are fully functional; global data streams, dynamic flight routes, and geographic markers are also fully laid out on the Earth's surface.
Looking at the finer details, atmospheric scattering, cloud rendering, day and night lighting effects, and the entire UI interface are all fully implemented.
Next, we put both models on the same competition arena.
AI blogger "Xiangyang Qiaomu" first ran three small tasks continuously with DeepSeek V4 Pro, then submitted the same prompts for two of the tasks to Grok 4.6 to directly compare the final outputs.
The first task is to call 3 Skills to develop and deploy a website.
DeepSeek smoothly completed the entire workflow, and its overall performance was very stable judging from the page design and completion level.
The second task is to replicate 60 different design styles at once, and generate a full set of Bento cards for centralized display.
In this round, DeepSeek made clear distinctions in fonts, color schemes, and layouts, delivering a pretty good overall visual effect.
The last task is to generate a 3D brick-breaking game from scratch.
The game is not only ready to play directly, but also added 3D scenes, background music, and dynamic sound effects, with fully responsive controls and great playability.
After that, the same prompts were submitted to Grok 4.6.
In the 3D brick-breaking game round, the two models were almost tied.
The game generated by Grok also has great texture, with visual effects, scene completeness, and playability on par with DeepSeek V4 Pro.
In the 60 Bento design round, Grok 4.6 scored one point back.
Some of the pages it generated are more mature in layout, color matching, and visual hierarchy, with better-looking final effects.
A more intense showdown took place on the Flappy Bird game development task.
Developer Jun Song gave exactly the same prompt, asking DeepSeek V4 Pro and Grok 4.6 to "handcraft" a game from scratch respectively.
As a result, the two models took completely different development paths.
DeepSeek V4 Pro consumed over 20,000 tokens at a total cost of only 0.019 USD; Grok 4.6 only used around 5,000 tokens, but its total cost reached 0.03 USD.
But judging from the final output, DeepSeek clearly performed better in this round.
The game not only features distant mountain landscapes and layered clouds, but also pipes with gradient effects and a sense of volume. When the character passes through the pipes, a floating "+1" animation even pops up on the screen.
Nearly all details from scene hierarchy to operation feedback are fully optimized, and the end-to-end construction completion level is significantly higher than that of Grok 4.6.
In another front-end test conducted by developer Hamza, DeepSeek V4 Pro once again outperformed Grok 4.6.
The output delivered by V4 Pro is superior in both page completion level and final visual effect.
However, V4 Pro did not deliver perfect performance in every test.
In a set of pelican comparison tests, compared to the Flash version, V4 Pro delivered higher overall image completion level and a more aesthetic elephant shape.
The only problem is — the movement direction of the pelican was drawn completely reversed.
When we further increased the difficulty, asking DeepSeek V4 Pro, GPT-5.6 Sol, and Claude Opus 5 to generate the same cherry blossom tree using Three.js, the gap became more obvious.
V4 Pro is inferior to the other two top-tier models, no matter in the details of the trunk and branches, or in the lighting, depth of field, and overall atmosphere.