The latest voice foundation model from Google DeepMind, integrating extended chain-of-thought reasoning directly into live spoken dialogues.
Gemini 3.5 Flash
Mid-generation multimodal speed model with configurable thinking levels, allowing developers to balance latency and reasoning depth.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
Gemini 3.5 Flash integrates thinking budget controls into the Flash architecture. Developers can tune reasoning effort from 0 (instant response) to maximum (deep analytical thinking), matching task complexity on the fly.
Strengths & Considerations
- Configurable thinking effort levels
- Large 64k token output capacity
- Outstanding price-to-performance ratio
- Maximum thinking levels increase latency
Token & API Pricing
Similar & Alternative Models
Explore other frontier models from Google and comparable reasoning engines.
The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.
Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.
Universal multilingual speech engine supporting live simultaneous translation and speaker diarization across 100+ languages.