Frontier Model Intelligence

AI Models Directory

Comprehensive directory of 29 foundation models, LLMs, and multimodal engines from top research labs. Real pricing, benchmark scores, and architecture cards.

Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

Context :1M ctx
Pricing :$0.75 / 1M tok ($1.50 reg)
Type :multimodal-realtime
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

Context :1M ctx
Pricing :$0.60 / 1M tok (Universal Omni)
Type :multimodal-realtime
Google

High-performance multimodal foundation model featuring algorithmic reasoning enhancements and agentic long-form video understanding.

Context :1M ctx
Pricing :$0.50 / 1M tok (Agentic Video)
Type :multimodal-realtime
Google

Second-generation on-device physical AI foundation model incorporating tactile sensor feedback for delicate robotic manipulation.

Context :33K ctx
Pricing :On-Device Edge Embedded 2.0
Type :multimodal-realtime
Google

Cloud embodied foundation model capable of synthesizing novel robotic tool use and executing complex mechanical assemblies.

Context :524K ctx
Pricing :Enterprise Robotics Cloud 2.0
Type :multimodal-realtime
Google

High-speed, high-density lightweight model providing strong analytical accuracy for 8 cents per million tokens.

Context :1M ctx
Pricing :$0.08 / 1M tok
Type :text-to-text
Google

Iterative performance leap in the Flash family featuring major upgrades in code refactoring and autonomous agent tool loops.

Context :1M ctx
Pricing :$0.35 / 1M tok
Type :multimodal-realtime
Google

Sub-second, sub-cent visual generation model designed for high-scale gaming asset generation, thumbnails, and preview workflows.

Context :262K ctx
Pricing :$0.008 / image
Type :text-to-image
Google

Mid-generation multimodal speed model with configurable thinking levels, allowing developers to balance latency and reasoning depth.

Context :1M ctx
Pricing :$0.25 / 1M tok (with Thinking Levels)
Type :multimodal-realtime
Google

Extended Reasoning (ER) foundation model enabling fine-motor tool use and dynamic physics adaptation for autonomous robotic arms.

Context :262K ctx
Pricing :Enterprise Robotics Cloud
Type :multimodal-realtime
Google

The Gemini 3.1 family’s budget workhorse, delivering solid reasoning and vision parsing at 6 cents per million tokens.

Context :1M ctx
Pricing :$0.06 / 1M tok
Type :text-to-text
Google

High-speed image generation and editing model engineered for conversational chats and real-time interactive design.

Context :524K ctx
Pricing :$0.02 / image
Type :text-to-image
Google

Google’s most advanced model for complex coding, scientific discovery, and long-horizon multimodal reasoning with a 64K output token buffer.

Context :1M ctx
Pricing :$1.50 / 1M tok (64k output)
Type :text-to-text
Google

The third-generation speed-optimized foundation model combining sub-second latency with near-Pro intelligence.

Context :1M ctx
Pricing :$0.15 / 1M tok
Type :multimodal-realtime
Google

Professional-grade visual synthesis model featuring flawless text rendering, spatial consistency, and multi-turn iterative image editing.

Context :524K ctx
Pricing :$0.04 / image (HD)
Type :text-to-image
Google

The inaugural model in the Gemini 3 family, establishing a new frontier in native multimodal reasoning and autonomous coding.

Context :1M ctx
Pricing :$1.50 / 1M tok
Type :text-to-text
Google

Specialized desktop and browser agent model that interprets screen pixels, executes mouse clicks, keyboard strokes, and multi-app tasks.

Context :524K ctx
Pricing :$1.00 / 1M tok (OS Agent)
Type :multimodal-realtime
Google

The most cost-effective tier in the Gemini 2.5 generation, providing high-volume text and vision comprehension for $0.05/1M tokens.

Context :1M ctx
Pricing :$0.05 / 1M tok
Type :text-to-text
Google

Cloud-scale spatial reasoning model providing high-level semantic navigation and multi-step plan generation for robotics fleets.

Context :131K ctx
Pricing :Enterprise Robotics Cloud
Type :multimodal-realtime
Google

Google’s dedicated reasoning model utilizing test-time compute to solve Olympiad-level mathematics and competitive coding.

Context :1M ctx
Pricing :$2.00 / 1M tok (with Thinking)
Type :reasoning-llm
Google

On-device vision-language-action foundation model for real-time robotic perception and physical manipulation.

Context :16K ctx
Pricing :On-Device Edge Embedded
Type :multimodal-realtime
Google

Google’s frontier reasoning and coding model with 2M token context, designed for complex analytical and mathematical workflows.

Context :2M ctx
Pricing :$1.25 / 1M tok
Type :text-to-text
Google

Google’s multimodal flagship speed model featuring real-time conversational audio and video streaming with sub-second response times.

Context :1M ctx
Pricing :$0.10 / 1M tok | $0.002 / min audio
Type :multimodal-realtime
Google

Ultra-low-cost, high-throughput model built for massive batch processing, summarization, and cost-sensitive applications.

Context :1M ctx
Pricing :$0.075 / 1M tok
Type :text-to-text
Google

Google’s breakthrough long-context model featuring an unprecedented 2-million-token window with Mixture of Experts efficiency.

Context :2M ctx
Pricing :$1.25 / 1M tok (1.5 Pro) | $0.075 (Flash)
Type :text-to-text