Elon Musk released the new Grok model, which topped the intelligent agent evaluation rankings and is half the price of OpenAI's models.
Report from Zhidx on July 30: Today, SpaceXAI under Elon Musk announced the launch of the new-generation speech model Grok Voice Think Fast 2.0, which is the most capable speech-to-speech model the company has developed to date.
Elon Musk posted two consecutive posts on X, with the first one announcing that "Grok Voice now ranks first in agent performance", and the second one directly calling on netizens to "try the new Grok Voice".
Posts released by Elon Musk
In the Speech to Speech Index of the authoritative evaluation institution Artificial Analysis, the high-inference-capability version of Grok Voice Think Fast 2.0 ranks second with a score of 82.9%, only behind Qwen Audio 3.0 Realtime Plus (84.1%); in the Tau Voice benchmark test measuring agent performance, this model ranks first with a score of 56.5%, beating speech models from manufacturers including OpenAI, Google, and Alibaba Qwen.
More notably, the model has an average first audio response time of only 0.70 seconds, making it the only model in the top 5 of the list with a response time lower than 1 second, which is a significant improvement from the 1.25 seconds of the previous generation Grok Voice Think Fast 1.0.
In terms of pricing, Grok Voice Think Fast 2.0 costs $0.08 per minute (equivalent to $4.80 per hour of input audio), which is higher than the $3.00 price of the previous generation, but is only about 45% of the price of GPT-Realtime-2.1 High.
In the comment section, many netizens have shown high expectations for the new model.
Some netizens said they hope to use Grok Voice to generate multi-character audio dramas.
Some other developers said that they had been mainly using Gemini before, and now they finally have a new alternative.
01. Many developers gave positive feedback after actual tests: Extremely fast response, more natural conversations
Nick White, CEO of AI company Helionova AI, shared his actual test experience on X.
He said he has completed the full-stack real-time speech upgrade for grok-voice-think-fast-2.0 on his product Tradecraft, calling it "absolutely a leap-forward progress".
In the actual test audio released by Nick White, he asked Grok Voice to play a customer who was angry about a missed appointment by maintenance staff and filed a complaint, while he himself played the customer service operator, and the two sides had a multi-turn real-time phone conversation around the late appointment.
During the whole process, the model hardly had obvious pauses, could continue to advance the plot according to the conversation content, and give real-time feedback to the customer service's replies.
Another developer integrated Think Fast 2.0 into the terminal tool GrokTerm to demonstrate the conversation speed of Grok Voice Think Fast 2.0.
In the video, the tester first asked the model to introduce the highlights of the new version, and the model responded at an extremely fast speed. Then the tester continuously issued instructions such as "switch the application theme to cyberpunk style", "switch back to the original theme", "switch to the other screen", "switch back to cyberpunk", the model responded quickly and completed all operations, answering fluently throughout the whole process.
Parker Conrad, a software engineer at SpaceXAI, also posted: "The entire speech team has poured a lot of effort, thinking and passion into this. The model's intelligence level, accuracy and capabilities are almost surreal. The speech field is still full of frictions, and Think Fast 2.0 has taken a big step towards the goal of 'out-of-the-box usability' — which is exactly what a speech model should be like."
02. Only 0.7 seconds response, faster inference and more accurate recognition
Positioned as a next-generation speech model, Grok Think Fast 2.0 focuses its upgrades on three aspects: model capability, response speed and inference efficiency.
The first is model capability.
According to the data released by Artificial Analysis, Think Fast 2.0 has a comprehensive score of 82.9%, which is 7.2 percentage points higher than the 75.7% of the previous generation.
Among them, it scored 97.2% in the Big Bench Audio speech reasoning test; the Full Duplex Bench that measures the multi-person real-time speech interaction experience reached 95.1%, a significant improvement over the 77.8% of the previous generation; the Tau Voice speech agent test reached 56.5%, making it the current No.1 on this list.
At the same time, the new model is also the fastest-responding model in the top 5 of the current ranking.
Its average first speech response time is only 0.70 seconds, faster than GPT-Realtime-2 High (1.14 seconds), GPT-Realtime-2.1 High (1.21 seconds) and Think Fast 1.0 (1.25 seconds).
In addition to speed, Grok Think Fast 2.0 also focuses on optimizing speech recognition.
In tests covering 24 languages and thousands of phrases, the overall transcription accuracy of Think Fast 2.0 is about 1.4 times higher than that of Think Fast 1.0, and about 1.5 to 2 times higher than that of Deepgram Nova 3 and ElevenLabs Scribe v2.
In real scenarios with high background noise and phone compression, this advantage is further expanded, and the recognition error can be reduced by about 10 times.
Another change comes from inference efficiency.
Different from many speech models that need to complete inference first and then generate speech, the Think Fast series supports completing inference while speaking, so it can handle complex tasks while maintaining low latency.
Think Fast 2.0 further optimizes the inference token utilization efficiency, with the median inference token consumption being only about 40% of that of the previous generation, which means that tool calls can usually start executing before the model finishes its first sentence, thus further shortening the overall response time.
In addition to inference efficiency, the conversation mode has also been adjusted. The R&D team used a large amount of reinforcement learning data to make the model closer to the real human communication mode. Compared with previous versions, the new model uses more short sentences, asks only one question at a time, and tries to avoid lengthy and repetitive expressions.
From the user's perspective, the entire communication process will be more natural, and even if the model is completing complex process planning in the background, users will not obviously feel the waiting time.
03. Conclusion: The competition of speech models has entered the "agent era"
Over the past year, manufacturers including OpenAI, Google and Alibaba Qwen have continued to upgrade their real-time speech models, and the focus of competition has gradually shifted from speech recognition and speech synthesis to complex task processing, multi-turn conversation and tool calling capabilities.
Judging from this release, SpaceXAI has continued this direction.
On the one hand, Think Fast 2.0 has further improved its capabilities in speech reasoning, real-time interaction and speech recognition; on the other hand, through "inference while speaking", lower inference token consumption and faster tool calling speed, it hopes to enable speech agents to truly take charge of real business scenarios such as customer service, telephone assistants and sales.
Judging from the latest list of Artificial Analysis, Think Fast 2.0 has already ranked first in speech agent capability (Tau Voice) and is at the forefront of the comprehensive Speech-to-Speech list.
With the completion of the default model upgrade on August 5, its actual performance will be further verified in more developer and enterprise applications.
This article is from the WeChat official account "Zhidx" (ID: zhidxc om), written by Jiang Yu, edited by Li Shuiqing, and authorized for release by 36Kr.