Just now, Elon Musk released Grok 4.7, but this time he has massively overhyped it.
There are no holidays in the AI industry.
From OpenAI, Anthropic to xAI, all major players seem to have made a tacit agreement to rush out new products overnight and seize market share intensively. Before the official holiday even starts, the atmosphere of the foundational model showdown has already reached its peak.
Just now, SpaceXAI officially released Grok 4.7.
Official blog link: https://x.ai/news/grok-4-7
The official positions it as the most powerful model for programming and knowledge work to date. The new model has been integrated into Cursor, Grok Build and Grok API, and will also be available through third-party programming tools, model routing platforms and cloud services.
The release of Grok 4.7 is much later than Elon Musk's original announcement. After teasing the model since late July, he revised the timeline to "wait a few more weeks", then narrowed it down to "within ten days", and finally added that "the model still needs further tuning"...
After multiple delays, the release focus of Grok 4.7 is very clear. SpaceXAI did not put the main emphasis on chat experience or multimodal demonstration, but concentrated on highlighting long-duration execution, self-verification, professional knowledge work, and more competitive price-performance ratio.
Grok 4.7: Bringing Long Tasks to the Table
According to the official introduction, Grok 4.7 uses a brand new foundational model larger than Grok 4.6, with longer reinforcement learning time and more difficult training tasks, a significant portion of which take several hours to complete.
The model also enhances long-context management and result verification capabilities, and is natively trained for the Grok Bot operating environment to improve performance in dialogue tasks and general knowledge work.
Multiple official published results show that Grok 4.7 has made significant progress compared to the previous generation.
In CursorBench 4.0, Grok 4.7 xHigh scores 46.3%, higher than Grok 4.6 High's 40.4%; DeepSWE v1.1 increases from 65.2% to 71.0%;
Terminal Bench 4.0 rises from 20.3% to 38.0%. The score of the electrical engineering benchmark EEBench increases from 53.0% to 64.0%, and AA Briefcase v1.1 for long-duration office tasks also rises from 1546 points to 1657 points.
After horizontal comparison, the advantages of Grok 4.7 are mainly concentrated in pricing and some professional tasks, and its overall performance has not yet taken the full lead.
On CursorBench 4.0 and Terminal Bench 4.0, its score is lower than Fable 5.1 Max; on DeepSWE v1.1, it is slightly lower than GPT 5.6 Sol Max;
But it scores 19.6% in the Harvey Legal Agent Benchmark, significantly higher than several comparison models in the official table.
Artificial Analysis later gave it an Intelligence Index score of 46 points, stating that Grok 4.7 has brought SpaceXAI into the top four AI laboratories worldwide, and its Coding Agent Index performance also exceeds GPT 5.6 Sol.
Pricing has become the most direct competitive chip for Grok 4.7. Its API pricing is the same as Grok 4.6, at $2 per million input tokens and $6 per million output tokens.
https://docs.x.ai/developers/pricing
SpaceXAI also provides a fast version with double the output speed, and the price is correspondingly doubled.
In Elon Musk's words, Grok 4.7 is a powerful product that integrates intelligence, speed and low cost. SpaceXAI officials directly claim that it can run twice as fast as similar models at half the price.
Apart from programming, SpaceXAI is starting to compete for professional work scenarios.
The official specifically emphasizes the model's capabilities for documents and presentations, citing evaluations such as GDPval and AA Briefcase to prove that the model can undertake part of the work of professionals such as lawyers, nurses and financial analysts.
Corresponding to the product direction, Grok 4.7 targets long-process tasks that require reading a large amount of materials, continuously calling tools and checking results repeatedly.
An insurance claim case demonstrated by Box can better illustrate this direction.
Grok 4.7 verified a $2 million claim, insurance policy and related records in Box Agent, found $82,000 in duplicate invoices and $64,000 in omitted supplier deductions, and also identified a deductible issue that could make the calculation result differ by $143,000.
The system finally generates a review report with citations, while insurance coverage confirmation and payment decisions are still completed by claims staff.
Security capabilities are also listed as an important part of this upgrade.
SpaceXAI says Grok 4.7 uses a brand new security protection system, scoring 62.4% in the LatchBio biosecurity benchmark; in HackerBench v0.3 for high-risk cybersecurity tasks, only 3.3% of dangerous dual-use prompts pass, while false rejections for normal security research are minimized.
Some cybersecurity partners will also get invited red team capabilities for defense research.
Netizens are going crazy testing, but public opinion is split...
The first batch of tests after Grok 4.7's launch mainly focus on web pages, 3D scenes and visual programming, but public opinion is sharply polarized.
Eric Zakariasson showed a comparison of Grok 4.6 and 4.7 making a demo project of *Age of Empires II*.
Ashutosh Shrivastava asked the model to build a Blender structure according to the floor plan and export it as an interactive Three.js website. All the above tests demonstrate its ability to handle multi-step visual tasks.
In another set of tests, developer Diego Cabezas asked Grok 4.7 xHigh to simulate a one-way street traffic light and randomly entering vehicles with Python, saying that its details are the richest in similar tests.
AI/ML API compared Grok 4.7 with Claude Opus 5 using four independent Three.js space scenarios, stating that the former spent about $0.20 on completing the test, while the latter cost about $2.05, and Grok took less time in three of the scenarios.
However, developer Bhavy believes that Grok 4.7's performance on 3D and front-end tasks is below expectations, with main problems including stiff physical effects, unstable prompt understanding and excessively fast token consumption.
Harshith generated the Three.js model of the Airbus H145 helicopter in Grok Build, also believing that the result of 4.7 is inferior to that of 4.6, and the task took about 45 minutes.
In another SVG web animation test of a pelican riding a bicycle, Grok 4.7's handling of the picture structure and motion relationship was also widely criticized.
Some developers gave a more neutral judgment. Salio believes that Grok 4.7's actual performance exceeds expectations, with visual quality slightly lower than Astra but higher than Sol.