
Now loading...
Google has launched its latest speech-to-text model, Gemini 3.5 Transcribe, which enhances intelligent voice interactions by delivering highly accurate transcriptions. Unlike traditional speech recognition systems that face challenges in noisy environments or complex language, this new model effectively transforms raw audio into well-structured text.
Users of Google’s Gemini app, as well as Android devices, have already experienced the advantages of this advanced transcription model, enjoying enhanced voice functionalities such as Rambler on Android. Developers can now incorporate similar features into their applications through the Gemini 3.5 Transcribe, accessible via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
Gemini 3.5 Transcribe integrates efficiently into developer workflows, facilitating the creation of voice agents, real-time captioning applications, and post-call analytics systems. The model is available through two distinct APIs: a real-time streaming option that provides continuous, low-latency interactions and a pre-recorded audio processing feature that can transcribe meetings, call logs, and other recorded content with speaker identification and precise timestamps.
This model is engineered to accurately capture users’ natural speaking patterns, allowing for better comprehension of intent and improved recognition of custom terminologies. Noteworthy features include smart transcription that rectifies self-corrections and eliminates filler words, the ability to assign complex tasks to other Gemini models, and exceptional transcription accuracy. With a reported average Word Error Rate of 4.0% for streaming and 2.6% for non-streaming contexts, it excels in noisy environments, effectively recognizing elements such as postal codes and order numbers.
Furthermore, Gemini 3.5 Transcribe is capable of adapting to specialized jargon and unique spellings, supporting over 85 languages and accommodating various regional accents and dialects. It can also accurately attribute speech in recorded content, providing timestamps for up to three speakers, with experimental support for more than three.
