Introducing Gemini 2.5 Flash and Gemini 2.5 Flash ImageGoogleReleased September 26, 2025

Gemini 2.5 Flash and Gemini 2.5 Flash Image

Sub-second multimodal model with native support for generating and editing images directly inline with text responses.

multimodal-realtimeProprietary API$0.125 / 1M tok | $0.03 / imageContext: 1.049M (1,048,576 tokens)Arena ELO: 1335

Technical Specifications

Architecture Type
Interleaved Multimodal Generation Transformer
Total Parameters
Undisclosed
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
8.192K (8,192 tokens)
Knowledge Cutoff
June 2025
Supported Modalities
text, code, vision, audio, image-gen
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Math500
82.4
Gen Eval
0.84
Mmlu Pro
71.2
Humaneval
90.1
Chatbot Arena ELO
1335

Deep Architectural Overview

Gemini 2.5 Flash introduces interleaved text-and-image outputs. Rather than calling a secondary image generation model, Gemini 2.5 Flash natively outputs pixel representations alongside conversational text.

Strengths & Considerations

Core Strengths
  • Native inline image generation and editing
  • Sub-second multimodal latency
  • Upgraded coding and tool calling capabilities
Known Limitations
  • Image resolution maxes at 1024x1024 natively without upscaling

Token & API Pricing

Input Tokens (1M)$0.13
Output Tokens (1M)$0.50
Cached Input (1M)$0.0312
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)