Introducing Gemini 2.0 Flash-LiteGoogleReleased April 10, 2025

Gemini 2.0 Flash-Lite

Ultra-low-cost, high-throughput model built for massive batch processing, summarization, and cost-sensitive applications.

text-to-textProprietary API$0.075 / 1M tokContext: 1.049M (1,048,576 tokens)Arena ELO: 1250

Technical Specifications

Architecture Type
Efficient Dense Transformer
Total Parameters
Undisclosed (optimized lightweight)
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
8.192K (8,192 tokens)
Knowledge Cutoff
August 2024
Supported Modalities
text, code, vision, audio
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Mmlu
78.4
Math500
62.5
Humaneval
80.2
Chatbot Arena ELO
1250

Deep Architectural Overview

Gemini 2.0 Flash-Lite delivers rapid inference speeds and industry-leading affordability without sacrificing 1M token context length. It is tailored for high-volume enterprise pipelines like document triage, categorization, and embeddings-augmented retrieval.

Strengths & Considerations

Core Strengths
  • Extremely low price ($0.075 / 1M tokens)
  • Fast Time-to-First-Token (TTFT)
  • Maintains full 1M context window
Known Limitations
  • Lower reasoning ceiling compared to 2.0 Flash and 2.5 Pro

Token & API Pricing

Input Tokens (1M)$0.07
Output Tokens (1M)$0.30
Cached Input (1M)$0.0188
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)