Grok 4.6 is now in the game, and Elon Musk is still one Odyssey short.
Elon Musk wants Grok to finish producing a full-length *Odyssey* by the end of this year.
This sounds like a typical Musk-style bold goal: using AI to challenge Nolan, Hollywood, and a $250 million production budget. But what truly deserves attention is not whether Grok can generate thousands of shots in five months, but whether it already has the ability to complete a complex task end-to-end.
After the release of Grok 4.6, this question has a more concrete answer.
In the latest round of LLM competition, it has broken into the first tier of large American models: its capabilities are close to the flagship models of OpenAI and Anthropic, its programming performance even leads on some lists, while its API price is significantly lower. At the same time, its shortcomings in long-chain Agent tasks still exist — it can understand problems, build frameworks, and start quickly, but it is not stable enough to wrap up tasks properly.
This is exactly the microcosm of the current large model competition.
As the capability gap narrows and Tokens get cheaper and cheaper, what truly determines commercial value is no longer just "how smart the model is", but also whether a task can be completed stably in one go, and how much it actually costs to finish one task.
For Musk, who invests tens of billions of dollars in capital expenditure in AI every year, Grok cannot just rush to the top of the list once in a while. It must stay at the top table continuously, and truly translate computing power, data and distribution into revenue.
01. Grok 4.6 is at the table, but still can't produce *Odyssey*
Before Nolan's *Odyssey* was released, Musk had already criticized the film for several rounds. He complained about the casting, the adaptation, and accused Nolan of "profaning Homer" for the sake of the Oscars. After the film was released, he simply set a goal on X: by the end of this year, Grok Imagine will produce a complete *Odyssey* that "conforms to history and stays true to Homer's art".
The three-minute *Odyssey* short film Musk reposted is indeed well-made. The armor, coast, and character close-ups all have the texture of a movie trailer, but between a three-minute polished short and a 90-minute feature film, there are still thousands of shots, plus a group of AI actors to develop.
Taking advantage of the Grok 4.6 release, we did not test videos again, but gave it two questions more suitable for language models. The first is to see if it can understand Homer. The second is to see if it can build a production system that expands the three-minute sample into a 90-minute feature film.
For the first question, Grok gave a more reliable answer than Musk himself.
"Conforming to history" and "staying true to Homer" sound not difficult, but when you actually open *Odyssey*, you will find that these two dimensions are hard to achieve at the same time. After hundreds of years of oral transmission, the weapons of the Bronze Age, the customs of the Iron Age, and the poet's mythological imagination are all stacked together. You can research what kind of sword Odysseus should use and what kind of ship he took, but you cannot dig out Cyclops, goddesses and the underworld from historical sites.
Grok split the problem into four layers: Bronze Age archaeology, the mixed time layers in Homer's epics, the poetics of the work itself, and Nolan's modern adaptation.
The solution it gave is that archeology is responsible for restricting artifacts, costumes and geography, and Homer is responsible for restricting characters, narratives and poetics. It also conveniently corrected a widely spread mistake: "the face that launched a thousand ships" comes from Marlowe, not Homer.
This question shows that Grok 4.6 not only knows how to look up information, but also can disassemble a proposition with internal conflicts, and construct the strongest arguments for both sides of the dispute.
The second question is closer to real production work.
Grok built an *Odyssey* production calculator, making the finished film duration, average shot length, first-pass usable rate, complex shot proportion, number of concurrent tasks, manual review time, cost and construction period all adjustable parameters.
The film production system generated by Grok 4.6
According to the default parameters, it estimates that a 90-minute feature film requires about 208 working days. Even Grok itself thinks this is a bit risky after calculation, and finally gave a very "realistic" suggestion: make a 12 to 20-minute selection of Homer's works.
Of course, 208 days cannot be taken as an actual measurement result. The first-pass usable rate, continuity failure rate, and how many times each usable shot needs to be generated are all assumptions set by the model. The calculator proves that Grok can disassemble engineering, establish formulas and build a production system, but does not prove that this system can really deliver the finished film within 208 days.
Moreover, it points out the hardest part of AI film making. It is no longer rare to generate a beautiful shot now. The difficulty is to keep the faces, costumes, ships, spaces and narratives consistent across thousands of shots, while keeping the scrap, rework and manual review within the range that producers can afford.
Conclusions from external lists are roughly the same.
On the Artificial Analysis list, Grok 4.6 has a comprehensive intelligence index of 61 points, the same score as GPT-5.6 Sol, and only 1 point away from Claude Fable 5. On CursorBench 3.2, Grok 4.6 High got 69.9%, exceeding Sol's 67.2%. The cost for it to complete one task is about $2.34, while Sol's is about $5.69. The starting price of Grok 4.6's API is $2 per million input Tokens and $6 per million output Tokens — the price is really competitive.
Image source / SpaceX
But for long-chain tasks such as DeepSWE and TerminalBench, Grok scored 65.9% and 26% respectively. Sol scored 73% and 34.6%, and Fable 5 scored 70% and 34.1%.
Grok is already very good at understanding problems, launching projects and building frameworks, but when it comes to the stage of continuous execution, handling unexpected situations and closing tasks stably, it still fails easily.
Overall, Grok 4.6 has taken a seat at the first table of American large models, and has played a very aggressive cost-performance card in front of the flagship models of OpenAI and Anthropic.
It can read Homer, write storyboards, calculate construction periods, and build production systems. But Odysseus drifted for ten years before returning to Ithaca. It is still very difficult for Grok to successfully "return home" with thousands of continuous shots before the end of this year.
02. A new generation every 35 days, Musk's top priority is to ensure Grok does not fall behind
From the public release of Grok 4.5 on July 8 to the release of Grok 4.6 on August 12, there was only a 35-day interval. Musk also previewed that Grok 4.7 with 2.1 trillion parameters will arrive in a few weeks, with stronger capabilities and higher Token efficiency, though slightly slower in reasoning.
4.7 has not been released yet, and its parameters and capabilities are only claimed by Musk for now, but 4.5 to 4.6 has at least completed an effective monthly upgrade. If 4.7 is launched on schedule, SpaceX AI will further demonstrate its ability to stay in the first tier.
This was somewhat unimaginable two years ago.
In November 2023, when Grok-1 was released, OpenAI had already launched GPT-4, Anthropic had already released Claude, and Google had TPU, search and cloud behind it. At that time, Grok was more like a chatbot included as a bonus for X Premium subscribers, with no enterprise customers, developer ecosystem or model transparency established, and its B-end market share was very low. In April this year, among the enterprise paid samples of Ramp customers, xAI's adoption rate was only 1.9%, while the figures for Anthropic and OpenAI were 34.4% and 32.3% respectively.
In the eyes of many practitioners, just like Meta, SpaceX AI's large models have long been stuck in the second tier.
But now Grok has squeezed over with a stool. The main reason it can catch up is that brute force brings miracles, relying on the combined hard push of computing power, data and engineering speed.
xAI built the first phase of Colossus with 100,000 Hopper GPUs in 122 days, and then expanded it to 200,000 units in 92 days. 19 days after the servers arrived on site, the cluster had already started running training tasks. Jensen Huang called this deployment speed "superhuman level". X provides Grok with real-time information, users and distribution entrances. Cursor provides real programming tasks and interaction data between developers and Agents. Cursor stated that the training of Grok 4.5 used trillions of Tokens of Cursor data.
The model does not need to be retrained from scratch every month. Grok 4.6 is built on the basis of 4.5, with main upgrades to supervised fine-tuning, reinforcement learning and Agent training environment. Colossus is responsible for scaling up the training scale, X and Cursor are responsible for continuously feeding back real tasks, and SpaceX AI uses post-training to integrate these tasks into the next version of the model.
After this assembly line runs smoothly, the version number can be pushed forward month by month.
Grok also has to run this fast, because SpaceX has bet too much money on AI.
Image source / pixabay
SpaceX's prospectus shows that in 2025, the AI division generated $3.2 billion in revenue, with an operating loss of $6.4 billion and capital expenditure of $12.7 billion, accounting for about 61% of the company's total capital expenditure. In the first quarter of 2026, SpaceX invested another $7.7 billion in the AI division, accounting for 76% of the company's total capital expenditure. In the second quarter of 2026, the capital expenditure of the AI division further rose to $15.828 billion, accounting for about 86% of the capital expenditure of the three major divisions, exceeding the full-year figure of 2025 in a single quarter; the cumulative amount in the first half of the year reached $23.6 billion.
The story SpaceX tells investors is already very clear. The company lists AI, Space and Connectivity as its three major business divisions, and calls AI the basic platform for future vertical integration. In the $28.5 trillion quantifiable market estimated by SpaceX itself, AI accounts for $26.5 trillion, while the rocket and satellite connectivity businesses add up to only about $2 trillion. Musk expects AI revenue to exceed the sum of SpaceX's other businesses in September 2026, and AI may contribute 99% of SpaceX's corporate value four or five years later.
Grok is the key variable for whether this heavy investment can get higher returns. The AI division can directly sell computing power to third parties, but only when its own model stays stably in the first tier can SpaceX have the opportunity to extend the same infrastructure upward to higher value-added businesses such as APIs, subscriptions and Agents.
If Grok stays in the first tier, SpaceX can make money from models, applications and computing power at the same time, and can also use its own model to create demand for Colossus and even future orbital computing power. If Grok falls behind, Colossus can still rent GPUs to Anthropic and Google, but SpaceX will regress from a full-stack AI company to a heavy-asset computing power landlord, responsible for building computer rooms, buying chips and paying electricity bills.
Moreover, with $23.6 billion in capital expenditure, if only a second-tier model is finally trained, Musk would probably be embarrassed to tell this story himself.
03. From competing on unit Token price to competing on "one-time successful delivery"
In the past few months, the competition for large models has been extremely fierce.
Anthropic relied on Claude Code and Fable 5 to build advantages in programming scenarios, OpenAI then used Codex and new models to catch up, and Grok updated two consecutive generations in more than a month.
The shelf life of top rankings has been shortened from several months to several weeks. When the capability gap between top models is only a few percentage points, another number will become particularly prominent — price.
Models from China have further pushed this pressure forward.
In July, Moonshot AI released Kimi K3 with 2.8 trillion parameters and 1 million Token context window. Its price is not particularly cheap, but opening the weights itself is weakening the top model vendors' control over inference prices.
DeepSeek went even further. As of August 13, the prices of DeepSeek V4 Flash per million input and output Tokens are only $0.14 and $0.28 respectively. According to the results announced by DeepSeek, its reasoning capability is already close to V4 Pro, and it can achieve similar performance for simple Agent tasks. On the same day, although DeepSeek announced that it would switch to peak-valley pricing from 00:00 Beijing time on August 17, even at peak prices, it is still significantly lower than Grok 4.6's $2 and $6.
OpenAI has also started to recalculate this account.
Three weeks after the release of GPT-5.6 on July 30, OpenAI cut the price of Luna by 80% and Terra by 20%. At present, Luna is $0.2 per million input Tokens and $1.2 per million output Tokens, while Terra is $2 and $12 respectively.
OpenAI's strategy is very clear: Sol handles the most complex judgment and planning, Terra takes on daily work, and Luna executes a large number of tasks with clear boundaries and high repetition. What it is striving for is to bring down the average cost of the entire workflow through model division of labor.