Moments ago, SpaceX launched its most powerful AI Grok 4.7, with an incredibly aggressive pricing that shocks the whole market, but its benchmark scores are disappointing.
Early this morning, SpaceXAI's most powerful model Grok 4.7 was released.
The release page opens with a highly "aggressive" positioning: the most powerful model for programming and knowledge work, with twice the speed of comparable models and half the price.
This upgrade is not limited to benchmark scores. Grok 4.7 adopts a larger base model, enhances capabilities in long-duration tasks, self-inspection and long-context management, and prioritizes testing of professional work including documents, presentations, legal, medical and engineering scenarios. Its target scenario is very clear: to allow the model to avoid detours in tasks lasting several hours and deliver results that can be used directly.
Elon Musk posted on X that "Grok 4.7 achieves a very competitive balance between intelligence level, operation speed and cost."
Long-duration tasks become the main battlefield of this generation of models
Grok 4.7 uses a larger base model than Grok 4.6. During the training phase, the reinforcement learning cycle is extended, and the task combination is further made more difficult, with more samples requiring several hours to complete. SpaceXAI claims that the new model has improved in checking its own work and managing long contexts.
Such improvements directly address the most intractable problems of current programming agents. Short code generation often only requires local judgment, while truly complex software tasks include reading repositories, breaking down requirements, modifying multiple files, running tests and repeated error correction. Whether the model can maintain its goals and detect errors in the long execution chain is usually more important than writing perfect code at one go.
CursorBench 4.0 is specifically designed to assess long-duration coding tasks. According to official data, Grok 4.7 scores 46.3%, while Grok 4.6 scores 40.4%, an increase of 5.9 percentage points. On DeepSWE v1.1, Grok 4.7's score at high inference intensity reaches 71.0%; its score on Terminal-Bench 4.0 rises from 20.3% of the previous generation to 38.0%.
Accordingly, Grok 4.7 is regarded as a model at the forefront in both price and performance on CursorBench 4.0.
The model is also trained to natively understand the operation framework of Grok Bot. According to the official statement, this adjustment improves the performance of conversational tasks and general knowledge work. In other words, the upgrade focus of Grok 4.7 has extended from single-turn responses to continuous collaboration between the model, tools and execution environments.
From writing code to full-fledged knowledge work
Programming is still the most prominent label of Grok 4.7, but the evaluation scope given by SpaceXAI is significantly broader. AA Briefcase and GDPval focus on multi-step work completed daily by professionals, covering common tasks for positions such as lawyers, nurses and financial analysts, as well as document and presentation production.
On AA Briefcase v1.1, Grok 4.7 scores 1657 points, higher than Grok 4.6's 1546 points; its GDPval Elo score rises from 1605 to 1695. In the official chart, its GDPval score is close to Fable 5.1's 1735 points, higher than GPT-6 Astra's 1542 points. SpaceXAI concludes that Grok 4.7 outperforms the previous generation in both tests, and its overall performance is comparable to other cutting-edge models.
Improvements in professional fields also appear in multiple directions. The EEBench score rises from 53.0% to 64.0%, the Harvey Legal Agent Benchmark score rises from 15.8% to 19.6%, and the HealthBench Professional score rises from 48.5% to 56.7%. These results come from the internal comparison table released by SpaceXAI, which can illustrate the changes between the old and new generations of models; cross-model comparison is still affected by inference intensity, tool configuration and test environment.
This set of results reveals a clear trend: the competition for cutting-edge models is entering the stage of "completing the entire work".
The model needs to understand tasks, call tools, output files and review them by itself, and the single question-and-answer capability is only one part of it. Grok 4.7 extends the duration of training tasks and strengthens self-verification, which is exactly to make up for this type of workflow.
Safety and Pricing
In addition to capability improvements, Grok 4.7 enables a brand new safety protection system. SpaceXAI states that it is the model with the best performance in refusing inappropriate responses and resisting jailbreak attempts among all models it has tested.
In terms of biosafety evaluation, Grok 4.7 tops the LatchBio benchmark with a score of 62.4%. The cybersecurity test HackerBench v0.3 mainly covers high-risk and malicious tasks. According to official data, the model only passes 3.3% of high-risk dual-use prompts, while rarely blocking legitimate security research requests by mistake.
SpaceXAI has also begun to invite a small number of cybersecurity partners to access its invitation-only red team capability for defense research. At present, this red team capability is provided in a controlled cooperation mode first.
The standard version retains the original price, and the speed of the fast version is doubled
Grok 4.7 has been launched on Cursor and Grok Build, and is available via Grok API, third-party programming frameworks, model routing services and cloud platforms. The starting price of the standard version is $2 per million input tokens and $6 per million output tokens, which is consistent with Grok 4.6.
This pricing strategy makes the upgrade even more compelling. Developers do not need to pay extra token costs for the standard version after the generation update, but can obtain stronger long-task performance, self-inspection capabilities and professional work results.
For programming agents and knowledge work products, while the unit price per model call is important, the number of attempts and reworks required to complete a task ultimately determines the actual cost.
To sum up this upgrade:
From the training method to the evaluation portfolio, Grok 4.7 is focusing on one thing: staying on the task for a long time and completing the complex work. Programming is only the first mature entry point, and documents, presentations, legal analysis, clinical reasoning and engineering tasks have all been included in the same capability map.
But netizens don't seem to buy it: "This model dares to be released at all, it is so bad that it should never have been launched." It can only be said that the selling point of Grok 4.7 this time is not to completely outperform competing products across the board. It uses API costs far lower than some flagship models to achieve performance that is very close to those models and leading in some individual professional tasks.
Reference links:
https://x.ai/news/grok-4-7
https://x.com/ArtificialAnlys/status/2102074904560771365
This article is from the WeChat official account Synced, edited by the editorial department of Synced, and published by 36Kr with authorization.