Google has officially introduced Gemini 3.5 Transcribe, a major new artificial intelligence model designed to transform how voice input is processed across real-time interfaces and pre-recorded audio archives. Arriving as part of the broader 3.5 branch expansion and succeeding the previous Chirp 3 speech-to-text engine, the new technology addresses long-standing frustrations with voice dictation by actively filtering out verbal stumbles, eliminating filler words, and outputting polished text. As conversational user interfaces and voice-driven developer tools become central to modern computing, this release marks a significant practical shift toward seamless audio integration across consumer devices and enterprise platforms.
Architecture and Performance Benchmarks of Gemini 3.5 Transcribe
The technological leap behind Gemini 3.5 Transcribe is rooted in significant gains in both processing speed and transcription accuracy. According to data cited by Google from independent benchmarking firm Artificial Analysis, the model achieves an average Word Error Rate of 4.0 percent for streaming applications and an even lower 2.6 percent for non-streaming audio processing. When evaluated on the rigorous FLEURS benchmark, the model records a Word Error Rate of 5.50 percent in streaming evaluations and 5.04 percent for non-streaming tasks. Furthermore, Google reports that the elapsed time from initial voice input to final transcribed text is roughly 70 percent faster than its predecessor, with live-speech error rates dropping to 5.5 percent.
Designed to accommodate a diverse global user base, the model natively supports more than 85 languages and regional dialects. It incorporates advanced linguistic capabilities such as automatic language detection, mid-sentence code-switching, and the identification of up to three distinct speakers in pre-recorded files, complete with word-level timestamps. Developers can also leverage custom vocabulary biasing to ensure that specialised industry jargon, acronyms, and proper nouns are transcribed correctly.
Ecosystem Integration and Dual API Deployment Options
To facilitate integration across diverse software environments, Google has deployed the technology through two distinct developer endpoints, each tailored to specific operational requirements. The Live API is engineered for continuous bidirectional streaming with sub-second latency, making it ideal for real-time voice agents and interactive assistants. Conversely, the Interactions API is optimised for longer audio files, meeting recordings, and call log analysis, providing robust speaker attribution and custom vocabulary support for up to 1,000 specific terms. Within Google's own ecosystem, the model is already powering features such as the Gboard Rambler tool on Pixel 11 devices in select regions, operating within Google Antigravity, supporting voice commands in the Gemini macOS app, and enabling direct dictation into web fields inside the Chrome browser.



