The latest voice foundation model from Google DeepMind, integrating extended chain-of-thought reasoning directly into live spoken dialogues.
AI Models Directory
Comprehensive directory of 29 foundation models, LLMs, and multimodal engines from top research labs. Real pricing, benchmark scores, and architecture cards.
The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.
Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.
Universal multilingual speech engine supporting live simultaneous translation and speaker diarization across 100+ languages.
High-performance multimodal foundation model featuring algorithmic reasoning enhancements and agentic long-form video understanding.
Second-generation on-device physical AI foundation model incorporating tactile sensor feedback for delicate robotic manipulation.
Cloud embodied foundation model capable of synthesizing novel robotic tool use and executing complex mechanical assemblies.
High-speed, high-density lightweight model providing strong analytical accuracy for 8 cents per million tokens.
Iterative performance leap in the Flash family featuring major upgrades in code refactoring and autonomous agent tool loops.
Sub-second, sub-cent visual generation model designed for high-scale gaming asset generation, thumbnails, and preview workflows.
Mid-generation multimodal speed model with configurable thinking levels, allowing developers to balance latency and reasoning depth.
Extended Reasoning (ER) foundation model enabling fine-motor tool use and dynamic physics adaptation for autonomous robotic arms.
Dedicated conversational voice model providing native bidirectional speech streaming with 220ms end-to-end latency.
The Gemini 3.1 family’s budget workhorse, delivering solid reasoning and vision parsing at 6 cents per million tokens.
High-speed image generation and editing model engineered for conversational chats and real-time interactive design.
Google’s most advanced model for complex coding, scientific discovery, and long-horizon multimodal reasoning with a 64K output token buffer.
The third-generation speed-optimized foundation model combining sub-second latency with near-Pro intelligence.
Professional-grade visual synthesis model featuring flawless text rendering, spatial consistency, and multi-turn iterative image editing.
The inaugural model in the Gemini 3 family, establishing a new frontier in native multimodal reasoning and autonomous coding.
Specialized desktop and browser agent model that interprets screen pixels, executes mouse clicks, keyboard strokes, and multi-app tasks.
Sub-second multimodal model with native support for generating and editing images directly inline with text responses.
The most cost-effective tier in the Gemini 2.5 generation, providing high-volume text and vision comprehension for $0.05/1M tokens.
Cloud-scale spatial reasoning model providing high-level semantic navigation and multi-step plan generation for robotics fleets.
Google’s dedicated reasoning model utilizing test-time compute to solve Olympiad-level mathematics and competitive coding.
On-device vision-language-action foundation model for real-time robotic perception and physical manipulation.
Google’s frontier reasoning and coding model with 2M token context, designed for complex analytical and mathematical workflows.
Google’s multimodal flagship speed model featuring real-time conversational audio and video streaming with sub-second response times.
Ultra-low-cost, high-throughput model built for massive batch processing, summarization, and cost-sensitive applications.
Google’s breakthrough long-context model featuring an unprecedented 2-million-token window with Mixture of Experts efficiency.