Introducing Gemini Omni FlashGoogleReleased August 27, 2026

Gemini Omni Flash

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

multimodal-realtimeProprietary API$0.60 / 1M tok (Universal Omni)Context: 1.049M (1,048,576 tokens)Arena ELO: 1450

Technical Specifications

Architecture Type
Universal Omnimodal Generation Transformer
Total Parameters
Undisclosed
Context Window
1.049M (1,048,576 tokens)
Max Output Tokens
32.768K (32,768 tokens)
Knowledge Cutoff
June 2026
Supported Modalities
text, code, vision, audio, video, 3d
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Omni Bench
87.2
Cross Modal Coherence
94.8
Chatbot Arena ELO
1450

Deep Architectural Overview

Gemini Omni Flash breaks down modal barriers entirely. It can take a voice command and video clip, edit the video, generate sound effects, narrate an explanation, and produce code in a single coherent inference stream.

Strengths & Considerations

Core Strengths
  • Universal any-to-any multimodal translation
  • Synchronized video and audio generation
  • Single unified API endpoint for all media modalities
Known Limitations
  • Video generation token consumption is higher than text queries

Token & API Pricing

Input Tokens (1M)$0.60
Output Tokens (1M)$2.50
Cached Input (1M)$0.1500
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

High-performance multimodal foundation model featuring algorithmic reasoning enhancements and agentic long-form video understanding.

1.049M ctx$0.50 / 1M tok (Agentic Video)