Elon Musk aims to squeeze into the "Big Three" with Grok 4.6
Who dares count how many posts Musk has made on X in the past two days?
Yesterday's theme was Grok Bot, and today it's Grok 4.6.
He is fully determined to break into the Top 3 — no, he is set on being Number 1.
Only 35 days have passed since the previous generation model Grok 4.5 was released, and Musk has already announced that Grok 4.7 will outperform all current models.
Let's hold off on drawing conclusions for now and take a closer look at how Grok 4.6 actually performs.
xAI's official positioning of Grok 4.6 is a model that is competent enough to "research a topic, analyze information, collaborate across code repositories, and turn an idea into a runnable application."
Models that support multi-step task execution are tailored for real-world work scenarios that require multiple rounds of reasoning, state retention, and tool coordination.
Challenges Grok 4.6 Aims to Solve
On the Artificial Analysis Intelligence Index, Grok 4.6 scored 61 points, ranking third globally — right behind Claude Opus 5 (63) and Claude Fable 5 (62), tied with GPT-5.6 Sol, and followed closely by Kimi K3.
In terms of generational comparison, Grok 4.6 is 5 points higher than the previous generation 4.5, and a full 23 points higher than Grok 4.3 released more than a month ago. Artificial Analysis commented that xAI has returned to the cutting-edge tier, "only lagging behind Anthropic".
In the dimension of knowledge work, Grok 4.6 scored 1753 points, slightly ahead of Claude Fable 5's 1741 points and GPT-5.6 Sol's 1728 points, ranking first among peer competitors.
This indicates that the model delivers the best overall performance in complex knowledge tasks such as long document analysis, cross-domain Q&A, and multi-step reasoning.
In the field of legal and complex professional reasoning, data from Harvey LAB is even more striking: Grok 4.6 reached 15.8%, far ahead of Fable 5 Max's 11.3%, while GPT-5.6 only scored 2.5%.
This means Grok 4.6 takes the lead in in-depth understanding of professional fields and multi-step derivation capabilities.
In terms of multi-round tool calling, the τ³-Banking simulation of real customer service business scenarios combined with tool calling gave Grok 4.6 a second-place finish at 50.7%, second only to Qwen3.8 Max's 51.3%. Its AA-Briefcase Elo score for long-range knowledge work scenarios is 1577, which falls into the Fable 5 tier and lags behind the Opus 5 series.
Michael Truell, CEO of Cursor, gave his endorsement immediately after the release, stating that this model is "significantly stronger on difficult tasks and knowledge work, combining Opus-level intelligence with low cost and high speed."
Some developers used Grok 4.6's toolchain to complete a five-stage dynamic simulation of river evolution — from the early pristine ecological river, to water transport trade, industrial era transformation, modern waterfront design, and finally the ecological restoration stage.
Throughout the entire multi-step task, the model demonstrated extremely strong long-range state retention and toolchain scheduling capabilities, with no context drift or tool call interruptions in the middle.
How Capable Is Grok 4.6 in Coding After the Acquisition of Cursor
SpaceX completed the acquisition of Cursor for approximately 60 billion US dollars earlier this year, and for this model iteration, the data value of Cursor has been systematically injected into the training process for the first time.
Specifically, the SFT phase of Grok 4.6 uses trajectory data regenerated by Grok 4.5, covering multiple fields including STEM, software engineering, and knowledge work, and filters problematic training samples through a model self-check mechanism.
In other words, the massive amount of real developer debugging data accumulated by Cursor has now been formally incorporated into the training pipeline.
The effect is immediate: Grok 4.6's score on DeepSWE v1.1 (which evaluates the model's code generation and debugging capabilities in real software engineering tasks) jumped sharply from 54% in version 4.5 to 65.9%, an increase of nearly 12 percentage points; its APEX-Agents score (which tests the completion rate of multi-step Agent tasks) rose from 47.1% to 57.5%, up 10 percentage points.
When @matt_palmer used Grok 4.6 to build a 2D agent simulation for the game *The Office*, his most immediate impression was the speed. With a throughput of approximately 80 tokens per second and low latency, he commented that it "rivals Cursor's Composer mode".
As we all know, in Agent programming scenarios that require high-frequency iteration, response speed directly affects the work rhythm.
When it comes to pricing, Grok 4.6's API is priced at $2 per million input tokens and $6 per million output tokens. When prompts enter xAI's long context window (starting from 200,000 tokens), the price doubles to $4/$12, and all tokens in the entire request are billed at the higher rate.
Even so, the short-context pricing is still more than 60% lower than that of Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). Data from Artificial Analysis shows that Grok 4.6 has an average per-task cost of approximately $0.84, making it the model with the best ratio of comprehensive capability to per-task cost currently available.
On the Reddit /r/cursor subforum, a developer posted that he uses Grok 4.6's High Reasoning mode for early architecture planning in his daily high-intensity development work, then distributes tasks to lightweight Agents to complete, and feels that "the cost-performance ratio is so high that it feels almost unreal".
Another user said they directly switched over from Claude Opus 4.8, because "it completes the same tasks for my use case, but saves almost 70% of the cost".
In Agent programming that heavily relies on context caching, each request for Claude Opus with equivalent capabilities costs about $0.052, while DeepSeek V4 Pro costs about $0.000875. Grok 4.6 falls between the two, but has obvious advantages in response speed and comprehensive intelligence.
Some testers used Grok 4.6 and Claude Opus 5 respectively to process three 3D scenes from Musk's TERAFAB large factory — the external environment at dusk, the clean room production workshop, and the central operation hall with walking workers. All scenes were retrieved directly via API without any editing.
Both models successfully delivered runnable code on the first attempt, but the cost gap is significant: Grok 4.6 spent $0.38, while Claude Opus 5 spent $2.06 — at the same output quality, the cost differs by 5.4 times.
In addition, Grok also uses significantly fewer tokens in each scenario.
Can Grok Break Into the New Top Tier of Three Leading Models?
At the end of 2025, the three leading models widely recognized by the industry were from OpenAI, Anthropic, and Google.
Looking at the top of the Artificial Analysis overall ranking now: Claude Opus 5 (63), Claude Fable 5 (62), Grok 4.6 (61), GPT-5.6 Sol (61), Kimi K3 (approximately 59.7).
Google is not in the top of this list.
It has not released any cutting-edge models since Gemini 3.1 Pro in February. Its flagship Gemini 3.5 Pro was reported to be months behind schedule, and has only been tested internally by partners so far; the latest releases from Google are 3.6 Flash, 3.5 Flash-Lite, and a security-focused vertical model called Flash Cyber.
With the departure of DeepMind's CEO and the loss of its reinforcement learning team, SemiAnalysis believes that DeepMind is no longer a cutting-edge laboratory.
The third leading position is half vacant.
But can the newly arrived Grok hold its ground?
Speaking of Grok, we cannot ignore the issues of extreme statements, excessive praise for Musk, and image generation features being used to create non-consensual sexual content. For enterprises with strict compliance and responsible AI requirements, some of these issues will not disappear just because the model scores better.
A month ago, during local initialization and each task execution, Grok Build CLI would, in a completely hidden manner, forcibly package the user's entire local project repository in the background and transmit it to SpaceXAI's cloud servers.
Therefore, after the release of version 4.6, there are also concerns raised by users: will it secretly package and take away my code repository?
However, for all models positioned between Grok 4.6 and DeepSeek V4 Pro, today is probably the most uncomfortable day of the year.
After all, the old saying goes: you don't mind your friend living a hard life, but you can't stand seeing your friend drive a luxury SUV.
This article is from the WeChat official account "APPSO", author: APPSO who discovers tomorrow's products, published with authorization from 36Kr.