The latest voice foundation model from Google DeepMind, integrating extended chain-of-thought reasoning directly into live spoken dialogues.
Gemini 2.5 Flash-Lite
The most cost-effective tier in the Gemini 2.5 generation, providing high-volume text and vision comprehension for $0.05/1M tokens.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
Gemini 2.5 Flash-Lite cuts inference cost in half while maintaining 1M token context window and strong factual recall, perfect for web scraping analysis, customer support bots, and real-time classification.
Strengths & Considerations
- Unbeatable $0.05 per 1M input token price
- High throughput for high-concurrency workloads
- Maintains 1M context window
- No audio streaming support
Token & API Pricing
Similar & Alternative Models
Explore other frontier models from Google and comparable reasoning engines.
The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.
Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.
Universal multilingual speech engine supporting live simultaneous translation and speaker diarization across 100+ languages.