HomeArticle

Google has launched its "most accurate" speech-to-text model to date, moving from "verbatim recording" to "comprehension-driven transcription".

36氪的朋友们2026-08-27 10:53
The commercial logic behind the battle for voice access points

On August 26 local time, Google released its latest speech-to-text model Gemini 3.5 Transcribe, which not only features higher recognition accuracy, but also extends its capability boundary from verbatim transcription to understanding speakers' intentions and automatically organizing the output results, marking that speech recognition technology has taken a step forward to deeper language understanding.

Different from traditional speech recognition tools, Gemini 3.5 Transcribe can identify speakers' self-corrections, automatically delete filler words, and directly convert unstructured spoken language into formatted text.

Google calls it the "most accurate" speech-to-text model to date, and has integrated it into Gboard for Android and the Gemini app for Mac, while opening preview access to developers via the Gemini API.

This release was launched on the same day as Gemini 3.5 Live and Gemini 3.5 Live Experimental, which together form Google's new Gemini Audio series.

Analysts pointed out that with the accelerated penetration of voice entry in Google's high-frequency products such as Chrome and Gboard, voice interaction is expected to gradually replace keyboard input and become the main connection method between users and AI systems.

As of press time, Google's stock price fell 1.37% in intraday trading on Thursday.

01

From "Verbatim Transcription" to "Intention Understanding"

The core upgrade of Gemini 3.5 Transcribe lies in its contextual understanding of natural language. Analysis shows that the key of this release is not only to improve the recognition accuracy in the traditional sense, but also to enable the model to understand speakers' intentions during the transcription process and "edit" the original speech.

Specifically, when a user says "Let's meet on Tuesday — no, Wednesday", the model can recognize that this is a self-correction, rather than recording both statements.

At the same time, the model will automatically delete filler words such as "um" and "uh", and complete the arrangement of punctuation marks and text formats. This model can directly convert messy, unstructured oral content into formatted text.

In terms of language coverage, this model supports more than 85 languages, has automatic language recognition capability, and can handle scenarios where multiple languages are mixed. Google also allows users to add custom vocabulary to make the model more accurately identify professional terms, special spellings, and alphanumeric combinations such as order numbers and postal codes. For pre-recorded audio, the model also supports speaker recognition and word-by-word timestamps.

02

Dual Progress on Developer Interface and Product Implementation

At the level of capability opening for developers, Gemini 3.5 Transcribe has entered the preview stage through the Gemini API, supporting two processing scenarios: real-time voice stream and recorded audio files.

According to Google's official documentation, the maximum supported duration for file transcription is 1 hour, while real-time transcription is oriented to low-latency application scenarios. Developers can call functions such as low-latency transcription, custom vocabulary and intelligent formatting to embed speech recognition capabilities into their own applications.

On the product side, Gemini 3.5 Transcribe has been deployed to the Gboard Rambler voice input function on the Android platform and the Gemini app for Mac.

Google also plans to integrate it into the Chrome browser, so that users can directly input content via voice in the web text input box, covering scenarios such as replying to messages, composing posts, and giving instructions to Gemini.

03

The Business Logic Behind the Competition for Voice Entry

This release is part of Google's systematic promotion of the Gemini "Voice Entry" strategy.

Google is further extending Gemini's capabilities from text and images to scenarios such as real-time conversation, speech translation and voice input. Gemini 3.5 Transcribe, Gemini 3.5 Live and Gemini 3.5 Live Experimental together form its Gemini Audio series.

From a commercialization perspective, speech-to-text is evolving from a simple back-end basic capability to an important interaction entry for AI applications. As the model has the integrated capability of "listening — understanding — organizing — executing", the interaction method between users and AI is expected to shift from keyboard input to voice commands.

For Google, embedding Gemini 3.5 Transcribe into high-frequency products such as Chrome, Gboard, Docs and Gmail is expected to further expand the penetration depth of Gemini in daily productivity scenarios, and consolidate its competitive position in the emerging entry of voice interaction.

This article is from the WeChat Official Account "Wall Street CN Max", Author: Yang Chen, published with authorization from 36Kr.