Introducing Gemini 2.5 Flash-LiteGoogleReleased September 26, 2025

Gemini 2.5 Flash-Lite

The most cost-effective tier in the Gemini 2.5 generation, providing high-volume text and vision comprehension for $0.05/1M tokens.

text-to-textProprietary API$0.05 / 1M tokContext: 1.049M (1,048,576 tokens)Arena ELO: 1265

Technical Specifications

Architecture Type
Efficient Dense Transformer
Total Parameters
Undisclosed
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
8.192K (8,192 tokens)
Knowledge Cutoff
June 2025
Supported Modalities
text, code, vision
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Mmlu
81.3
Math500
69.4
Humaneval
83.5
Chatbot Arena ELO
1265

Deep Architectural Overview

Gemini 2.5 Flash-Lite cuts inference cost in half while maintaining 1M token context window and strong factual recall, perfect for web scraping analysis, customer support bots, and real-time classification.

Strengths & Considerations

Core Strengths
  • Unbeatable $0.05 per 1M input token price
  • High throughput for high-concurrency workloads
  • Maintains 1M context window
Known Limitations
  • No audio streaming support

Token & API Pricing

Input Tokens (1M)$0.05
Output Tokens (1M)$0.20
Cached Input (1M)$0.0125
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)