As low as $0.54 per hour, Google starts to "PS" audio.
Following its further investment in real-time voice interaction and speech transcription, Google has extended its product line to the field of audio generation.
On September 23 local time in the US, Google officially launched two text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Google defines them as highly expressive audio generation models, aiming to push speech generation to evolve from static preset timbres to sound design with deep control capabilities.
At present, the two models have been launched on the Gemini API and Google AI Studio, and integrated into Gemini Notebook and Google Vids, with enterprise-level API access also opened simultaneously.
The credited team for this release includes Alan Cowen.
Cowen was previously the founder of Hume AI, a tech startup focused on "Empathic AI", and joined Google DeepMind this year as Director of Research Science. This time, he and Leland Rechis, Group Product Manager, are jointly responsible for the R&D and release of this speech project.
Against the backdrop that speech companies such as ElevenLabs continue to expand their commercial markets, Google's this release improves the control capability of speech models, and also tries to push speech generation capabilities to more application scenarios via APIs and its existing product ecosystem.
01 From "Selecting Voices" to "Designing Voices"
Compared with previous TTS models that mainly control voice style and rhythm through text prompts and audio tags, the Gemini 3.8 series focuses on the control of voice identity design and line-by-line performance.
The two newly launched models have different focuses: the flagship version Gemini 3.8 Flash TTS is mainly oriented to in-depth creativity and character design, covering scenarios such as games, immersive audiobooks and podcasts; the lightweight version Gemini 3.8 Flash-Lite TTS is oriented to high concurrency and batch processing demands, focusing on long video dubbing, audio content creation and customer service voice agents.
In terms of capabilities, Flash TTS shows the following changes:
· First is Generative Voice Design. No pre-recording is required, users can describe the character features, age, accent and timbre through natural language (for example: "A tired Berlin native who speaks English with a strong German accent at 2 a.m."), and the model will generate the corresponding timbre.
· Second is Voice Replication. According to the official description, the model supports voice replication based on about 30 seconds of reference audio to extract timbre features.
At the same time, the model supports line-by-line script control and non-verbal expression. Users can add Stage Directions in the text script, and insert response words such as <laughs>, <sigh>, <gasp> and |mhm| to simulate pauses and tone fluctuations in human conversations.
The last is two-person dialogue and timbre consistency for long audio. Users can generate alternating dialogues between two people based on one script, and control the tone and rhythm of both sides; for long audio generation, Google also optimizes timbre consistency to reduce timbre drift during continuous generation.
Gemini 3.8 Flash TTS converts the uploaded script into voices of different characters
02 High Benchmark Scores, But Real Experience Still Has Limitations
According to the test results released by Google, Gemini 3.8 Flash TTS has achieved high scores in multiple speech generation evaluations.
In the Voice Design Benchmark test of Hume AI, Flash TTS got an overall score of 71.4, and a score of 60.8 in the accent modeling dimension; in the Hume AI Overall Quality Index, Flash TTS and Flash-Lite TTS ranked top two respectively.
In addition, in the double-blind preference test of Voice Arena, the two models achieved high win rates in the evaluation of multiple languages such as Japanese, Modern Standard Arabic, Mexican Spanish, and Hindi.
However, the benchmark test scores do not fully represent the actual production experience.
Unfiltered AI, a tech media outlet, and its affiliated industrial analysis platform pointed out in the evaluation that although the new model performs outstandingly in specific complex accents and dynamic emotion control, tiny synthetic noises can still be observed in some high-frequency audio; in the later stage of continuous rendering of extremely long texts, a small number of samples also show slight timbre drift.
Meanwhile, the Model Card released by Google also shows that the new model may still have problems common to generative audio models, such as deviation from the text or abnormal pronunciation, and response delay or timeout may also occur under high load.
03 Google Starts to Play the Price Card
In addition to the improvement of model control, pricing is also one of the key focuses of Google this time.
In terms of pricing strategy, before the end of 2026, developers can enjoy the following API call prices: the text input price of Flash TTS is $0.50 per million Tokens, and the audio output price is $9.00 per million Tokens; the text input price of Flash-Lite TTS is also $0.50 per million Tokens, while the audio output price is reduced to $6.00 per million Tokens.
According to the conversion method between Tokens and audio duration given by Google (1 second of generated audio is approximately equal to 25 audio Tokens), the pure output API cost for continuously generating 1 hour of audio is about $0.81 for Flash TTS and about $0.54 for Flash-Lite TTS (equivalent to about 3.8 RMB).
Even converted at the standard price restored in 2027 (Flash TTS will be restored to $18 per million Tokens, Flash-Lite TTS will be restored to $12 per million Tokens), the corresponding output cost for 1 hour of audio is also at a relatively low level.
However, the above costs only refer to the theoretical conversion of audio output Tokens, and do not include extra expenses such as retries.
04 The Logic Behind Google's "PS" for Voices
For developers and enterprise users, calls and deployments can be made through the Gemini API and Google AI Studio; for end applications and office scenarios, Flash TTS is integrated into Gemini Notebook (note-taking and knowledge base analysis tool), while Flash-Lite TTS is directly built into Google Vids (AI video creation tool).
In this way, the competition of TTS is no longer limited to model parameters and single-point sound quality, but extends to APIs, office applications and the entire AI product ecosystem.
However, as the sample duration required for voice replication is shortened to about 30 seconds, how to establish a more complete authorization mechanism, responsibility identification and copyright arbitration system is a long-term issue that the whole industry needs to face.
Google has also set up protection measures in this release.
In terms of informed consent mechanism, Google requires that before voice replication, the person whose voice is to be replicated must first record an oral authorization statement, and the system will compare the voiceprint of the authorized recording with the reference audio to confirm consistency before generation.
In terms of digital watermarking and traceability, all audio generated by Gemini Audio is embedded with SynthID audio watermarks that are imperceptible to the naked eye and human ear, and comes with C2PA digital credentials to provide traceability information of the generated content.
In addition, Google currently has restrictions on the available regions for the voice replication function, which is temporarily not open in regions such as Illinois and Texas in the United States, the European Economic Area (EEA), the United Kingdom, Switzerland and India.
Special contributor Wu Ji also contributed to this article
This article is from "Tencent Tech", Author: Su Yang, Editor: Xu Qingyang, 36Kr is authorized to republish.