Gemini 3.8 Live and 3.5 Transcribe Upend Voice App Development

The release of Google DeepMind's Gemini Audio models marks a significant advancement in the creation of real-time voice applications. With the Gemini 3.8 Live and its Enhanced Thinking variant, developers can harness real-time voice interaction across a spectrum of uses, including native speech-to-speech capabilities and task execution while maintaining dialogue integrity. This leap in functionality, emphasized by Google, positions the model as a frontrunner on the Artificial Analysis Speech-to-Speech leaderboard. Complementing this is Gemini 3.5 Transcribe, whose transcriptions boast commendable accuracy, evidenced by a Word Error Rate (WER) of 4.0% for streaming and an impressive 2.6% for non-streaming audio, as highlighted by Google Blog.
The versatility of these models is further enhanced by their multilingual support, spanning over 97 languages — a critical feature for global applications. Developers will find the Google AI Studio platform essential for building these cutting-edge AI applications, integrating seamlessly through the Live API. The models are priced accessibly, with audio input and output costs set at $0.005 per minute and $0.018 per minute, respectively. Additionally, developers can personalize their applications further with a custom vocabulary list that supports up to 1,000 terms.
The partners involved, including Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents, provide robust media streaming infrastructure that ensures the efficient deployment of these capabilities. These partnerships fortify the Gemini Audio models' integration into existing systems, allowing developers to craft voice applications with real-time speech understanding, an essential feature for voice-first interfaces.
In forthcoming applications, the deployment of Gemini models should lead to more robust performance in real-time dialogue and task-related scenarios. Smart transcription mode available with Gemini 3.5 Transcribe delivers polished transcripts by applying structured formatting, self-corrections, and eliminating disfluent elements like filler words. Such enhancements promise to elevate user experiences, ensuring interactions are not only efficient but also comprehendible and seamless, paving the way for more natural integration into everyday technology interactions.
Intelligence
Why this matters
- Real-time voice applications enhance user interaction capabilities.
- Multilingual support broadens market reach for developers.
- Affordable pricing encourages widespread adoption of voice technology.
Keep independent coverage alive.
No ads. No paywall. No corporate backing. Just sharp, weekly intelligence on the language industry — free, because it should be.