The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.
Gemini 3.8 Audio (Live, Live Extended Thinking)
The latest voice foundation model from Google DeepMind, integrating extended chain-of-thought reasoning directly into live spoken dialogues.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
Gemini 3.8 Audio (Live, Live Extended Thinking) allows users to speak directly with an AI that reasons deeply before answering verbally. It combines sub-200ms conversational responsiveness with high-level cognitive problem solving.
Strengths & Considerations
- Live conversational audio with extended reasoning compute
- 195ms conversational responsiveness
- Solves spoken complex math and engineering problems without text intermediation
- Extended thinking on difficult queries introduces purposeful thinking pauses
Token & API Pricing
Similar & Alternative Models
Explore other frontier models from Google and comparable reasoning engines.
Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.
Universal multilingual speech engine supporting live simultaneous translation and speaker diarization across 100+ languages.
High-performance multimodal foundation model featuring algorithmic reasoning enhancements and agentic long-form video understanding.