Introducing Gemini 3.5 Flash-LiteGoogleReleased July 21, 2026

Gemini 3.5 Flash-Lite

High-speed, high-density lightweight model providing strong analytical accuracy for 8 cents per million tokens.

text-to-textProprietary API$0.08 / 1M tokContext: 1.049M (1,048,576 tokens)Arena ELO: 1320

Technical Specifications

Architecture Type
Efficient Dense Transformer
Total Parameters
Undisclosed
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
16.384K (16,384 tokens)
Knowledge Cutoff
April 2026
Supported Modalities
text, code, vision
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Mmlu
85.1
Math500
79.4
Humaneval
89.2
Chatbot Arena ELO
1320

Deep Architectural Overview

Gemini 3.5 Flash-Lite represents the evolution of budget AI, achieving benchmarks that rival early Gemini Pro models while keeping costs virtually negligible.

Strengths & Considerations

Core Strengths
  • 8 cents per million tokens
  • 16k output token length
  • High throughput reliability for enterprise workloads
Known Limitations
  • Lacks native audio output

Token & API Pricing

Input Tokens (1M)$0.08
Output Tokens (1M)$0.32
Cached Input (1M)$0.0200
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)