Introducing Gemini 2.0 FlashGoogleReleased April 15, 2025

Gemini 2.0 Flash

Google’s multimodal flagship speed model featuring real-time conversational audio and video streaming with sub-second response times.

multimodal-realtimeProprietary API$0.10 / 1M tok | $0.002 / min audioContext: 1.049M (1,048,576 tokens)Arena ELO: 1315

Technical Specifications

Architecture Type
Multimodal Streaming Transformer
Total Parameters
Undisclosed
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
8.192K (8,192 tokens)
Knowledge Cutoff
August 2024
Supported Modalities
text, code, vision, audio, video
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Gpqa
53.4
Math500
76.5
Mmlu Pro
66.8
Humaneval
88.2
Chatbot Arena ELO
1315

Deep Architectural Overview

Gemini 2.0 Flash is a transformative multimodal workhorse. It powers the Multimodal Live API with native speech-to-speech interaction, bidirectional visual reasoning from camera feeds, and robust function calling across external tool ecosystems.

Strengths & Considerations

Core Strengths
  • Native Multimodal Live audio/video streaming
  • Fast sub-second response times
  • Top-tier tool calling and structured output accuracy
Known Limitations
  • Complex multi-step coding trails dedicated reasoning models

Token & API Pricing

Input Tokens (1M)$0.10
Output Tokens (1M)$0.40
Cached Input (1M)$0.0250
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)