The latest voice foundation model from Google DeepMind, integrating extended chain-of-thought reasoning directly into live spoken dialogues.
Gemini 3.5 Audio (Live Translate, Transcribe, Transcribe Live)
Universal multilingual speech engine supporting live simultaneous translation and speaker diarization across 100+ languages.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
Gemini 3.5 Audio is built for global communication, live broadcast subtitling, and international meeting transcription. It automatically identifies multiple simultaneous speakers and outputs translated speech in real time.
Strengths & Considerations
- Real-time simultaneous translation across 100+ languages
- 3.6% Multilingual Word Error Rate (WER)
- Advanced multi-speaker diarization in noisy environments
- Exclusively optimized for audio, speech, and translation workflows
Token & API Pricing
Similar & Alternative Models
Explore other frontier models from Google and comparable reasoning engines.
The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.
Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.
High-performance multimodal foundation model featuring algorithmic reasoning enhancements and agentic long-form video understanding.