Introducing Gemini 3.5 Audio (Live Translate, Transcribe, Transcribe Live)GoogleReleased August 26, 2026

Gemini 3.5 Audio (Live Translate, Transcribe, Transcribe Live)

Universal multilingual speech engine supporting live simultaneous translation and speaker diarization across 100+ languages.

multimodal-realtimeProprietary API$0.0025 / min audio (Translate & Live)Context: 262.144K (262,144 tokens)

Technical Specifications

Architecture Type
Massive Multilingual Speech Transformer
Total Parameters
Undisclosed
Context Window
262.144K (262,144 tokens)
Max Output Tokens
16.384K (16,384 tokens)
Knowledge Cutoff
June 2026
Supported Modalities
audio, text, speech, translation
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Multilingual Wer
3.6
Translation Bleu
44.8
Diarization Accuracy
97.1

Deep Architectural Overview

Gemini 3.5 Audio is built for global communication, live broadcast subtitling, and international meeting transcription. It automatically identifies multiple simultaneous speakers and outputs translated speech in real time.

Strengths & Considerations

Core Strengths
  • Real-time simultaneous translation across 100+ languages
  • 3.6% Multilingual Word Error Rate (WER)
  • Advanced multi-speaker diarization in noisy environments
Known Limitations
  • Exclusively optimized for audio, speech, and translation workflows

Token & API Pricing

Input Tokens (1M)$0.30
Output Tokens (1M)$1.20
Cached Input (1M)$0.0750
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)
Google

High-performance multimodal foundation model featuring algorithmic reasoning enhancements and agentic long-form video understanding.

1.049M ctx$0.50 / 1M tok (Agentic Video)