Introducing Gemini 3.1 Flash Audio (Flash Live, TTS)GoogleReleased April 15, 2026

Gemini 3.1 Flash Audio (Flash Live, TTS)

Dedicated conversational voice model providing native bidirectional speech streaming with 220ms end-to-end latency.

multimodal-realtimeProprietary API$0.003 / min audio (Live)Context: 131.072K (131,072 tokens)

Technical Specifications

Architecture Type
Native Speech-to-Speech Transformer
Total Parameters
Undisclosed
Context Window
131.072K (131,072 tokens)
Max Output Tokens
8.192K (8,192 tokens)
Knowledge Cutoff
December 2025
Supported Modalities
audio, text, speech
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Speech Wer
4.1
Audio Latency Ms
220
Emotion Fidelity
94.2

Deep Architectural Overview

Gemini 3.1 Flash Audio eliminates traditional text-to-speech transcoding pipelines. By processing and synthesizing acoustic tokens natively, it conveys subtle emotional inflection, natural pauses, and human-like interruption handling.

Strengths & Considerations

Core Strengths
  • 220ms glass-to-glass audio latency
  • Native emotional prosody and accent synthesis
  • Seamless interruption handling
Known Limitations
  • Context limited to 128k audio tokens

Token & API Pricing

Input Tokens (1M)$0.25
Output Tokens (1M)$1.00
Cached Input (1M)$0.0625
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)