Meta's flagship model has staged a dramatic comeback, which is even cheaper than DeepSeek, and its Chinese-American top executive is directly challenging Gemini head-on.
Zhidx reported on September 3 that Meta today released its new generation of top-tier flagship model Muse Spark 1.3. On the AI analysis index ranking of Artificial Analysis, the new model scores second only to Claude Fable 5.1 and Claude Opus 5.
Mark Zuckerberg, co-founder and CEO of Meta, stated on social platform X that the performance of Muse Spark 1.3 is beyond imagination.
Alexander Wang, head of Meta's Superintelligence Lab and Chief AI Officer, said the model has negligibly low cost, with greatly improved agent capabilities, programming capabilities and ease of use. Moreover, when Muse Spark 1.3 is paired with Muse Code, its evaluation performance is comparable to that of Claude Code + Opus 5 and Claude Code + Fable 5.
Some developers used Muse Spark 1.3 to create the Minecraft game at a total cost of only 10 cents (about 0.67 yuan).
Just 4 hours before the release of Meta's new model, Google launched its two versions of Gemini 3.8, the most powerful reasoning and programming model to date: Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. However, Meta soon overtook it on the Artificial Analysis ranking. Currently on the AI analysis index list, Muse Spark 1.3 scores 62 points, second only to Claude Fable 5.1 (66 points) and Claude Opus 5 (63 points), while Gemini 3.8 Flash scores 59 points.
Google only led on the Artificial Analysis index for a few hours before being overtaken by Meta. Alexander Wang also directly challenged Google on social platforms, posting "gemini who?" and adding that "It's not too late for Gemini to cancel (the launch) now."
In the past 5 months, Meta has successively launched four Muse Spark models: the debut Muse Spark in April, the updated Muse Spark 1.1 in July, Muse Spark 1.2 launched in August, plus the latest Muse Spark 1.3, its iteration rhythm is getting increasingly intensive.
As can be seen from the benchmark test chart released by Meta, Muse Spark 1.3 has won 5 first places. It outperforms GPT-5.6 Sol and Claude Opus 5 in most tests of long context and programming, and only draws with GPT-5.6 Sol in the Terminal-Bench terminal operation Agent test. Its relatively weak part is the Agent capability, with low scores in web search Agent scenarios and the general knowledge task GDPVal-AA v2.
In terms of pricing, the API pricing of Muse Spark 1.3 remains consistent with that of its predecessor 1.2: $1.25 per million input tokens (cache miss) (about 8.4 yuan), $0.15 per million input tokens (cache hit) (about 1 yuan), and $4.25 per million output tokens (about 28.55 yuan). The price of Muse Spark 1.3 is slightly higher than that of Gemini 3.8 Flash, and cheaper than the peak price of DeepSeek-V4-Pro.
Price comparison of some large models (chart by Zhidx)
Muse Spark 1.3 has been gradually rolled out on Muse Code and Meta Model API, with the previously supported reasoning modes now available; the max-reasoning (extreme reasoning mode) will be opened soon after completing additional security tests.
Meta said it has planned an exciting product roadmap, including models with larger parameter scales, the release of Muse Spark open-weight versions, and more updates.
01. Developer Field Test: Fast Speed, Low Cost, But Ordinary Generation Effect
On social platform X, many developers shared their actual test cases of Muse Spark 1.3, and most of them commented that it is fast and low-cost, but its generation performance is not amazing.
One developer compared the 3D simulation effects of Muse Spark 1.1 to 1.3 for the same building based only on photos, and found that there were significant improvements in geometric shape, structure and visual effect, while the total cost remained unchanged at 0.6 dollars (about 4.03 yuan).
Another developer shared the comparison effect between the Muse Spark 1.3 Ultra Contributor version (with Ultra ultra-high computing power mode enabled) and Claude Fable 5.1 xHigh, with the same prompt word, both running for about 2 hours. The prompt word has a built-in set of self-iteration optimization rules: if the independent evaluation model scores the output result below 9.5/10, the model must continue to optimize and re-attempt.
To his shock, Muse Spark kept iterating and ran a total of 20 rounds of self-optimization loops, with 3 agents started in each round; it performed more than 60 agent tasks in two hours, but the total cost was less than 1 dollar (about 6.7 yuan).
One netizen questioned in the comment section whether this means that Muse Spark 1.3 never completed the task and just stopped after spending 1 dollar (about 6.7 yuan).
In MineBench, the open-source 3D spatial reasoning benchmark test, the actual test comparison between Gemini 3.8 Flash and Muse Spark 1.3 shows that the total test cost of Gemini 3.8 Flash is 1.18 dollars (about 7.93 yuan), the average reasoning time is 1 minute and 51 seconds, and the first test Elo score is 1888; the total test cost of Muse Spark 1.3 is 6.57 dollars (about 44.14 yuan), the average reasoning time is 5 minutes and 36 seconds, and the first test score is 1787, which means there is still a certain gap between its output quality and that of the top cutting-edge models.
Another developer compared the effects of Muse Spark 1.3, Qwen 3.8 Max, GLM 5.3 and GPT-5.6 Sol, and his overall feeling is that Muse Spark 1.3 is comparable to Qwen 3.8 Max in terms of cost and speed, while GLM 5.3 has the best overall performance in performance, cost, speed and token consumption.
Although Muse Spark 1.3 has a very low price and fast running speed, its performance is only considered decent. Qwen 3.8 Max has the best output quality and highest completion degree, but it takes about an hour and a half to run, while Muse Spark only takes about 2 minutes.
02. No Deviation in Long-cycle Tasks, 25% Reduction in Programming Token Consumption
Meta's blog mentions that the advantages of Muse Spark 1.3 lie in supporting long-cycle tasks and programming scenarios.
Muse Spark 1.3 can collaborate with users to process multiple workflows in parallel in a single long conversation session. When facing open-ended goals, it will call tools to independently construct context from messy and contradictory information sources, then actively correct loopholes in the scheme, and output complete deliverables based on the obtained information.
At the same time, Meta's researchers completed model training based on multiple diverse evaluation frameworks, so that it can adapt to various agent operating environments.
For example, as shown in the figure below, the prompt word obtained by Muse Spark 1.3 is very complex: it needs to find event details, copywriting and pictures in different locations, then edit and combine them to create and publish on the ad management platform.
Compared with previous Muse Spark models, Muse Spark 1.3 can better retain detail requirements in multi-step tasks, will not miss constraints, and will not deviate from the established workflow.
Meta also improved the multi-task processing capability of the model. For example, even in a single session context with messy information, no matter whether the user continues a past task or interrupts the task midway, the model can match the new input prompt to the corresponding task.
Muse Spark 1.3 can understand what it can do and what it cannot do, what it knows and what it does not know, and how to respond when encountering obstacles, instead of producing wrong prediction results.
For example, when the prompt word is "You are a mechanical engineer at a small aerospace company, designing an experimental X-type wing component for the next-generation aircraft. To support the design review, please write a first draft of the fluid simulation report based on the attached materials: (1) Preliminary CFD simulation results; (2) The CAD model STEP file of the wing component used for this simulation.", Muse Spark 1.3 will generate the corresponding files:
In terms of programming, compared with Muse Spark 1.2, Muse Spark 1.3 requires fewer operation steps, has more concise expressions, and the overall coding style is clearer. According to the test results of Meta engineers, its number of tool calls is reduced by 20%, and the number of tokens used is reduced by 25%.