Introducing Gemini 2.5 Computer UseGoogleReleased October 7, 2025

Gemini 2.5 Computer Use

Specialized desktop and browser agent model that interprets screen pixels, executes mouse clicks, keyboard strokes, and multi-app tasks.

multimodal-realtimeProprietary API$1.00 / 1M tok (OS Agent)Context: 524.288K (524,288 tokens)

Technical Specifications

Architecture Type
Agentic Vision-Action Transformer
Total Parameters
Undisclosed
Context Window
524.288K (524,288 tokens)
Max Output Tokens
8.192K (8,192 tokens)
Knowledge Cutoff
June 2025
Supported Modalities
vision, text, screen-actions
License & Access
Google Cloud API Terms of Service

Benchmark Evaluations

Os World
44.8
Web Arena
52.3
Screen Click Accuracy
94.6

Deep Architectural Overview

Gemini 2.5 Computer Use provides native operating system automation. It perceives graphical user interfaces (GUIs), accurately locates interactive UI elements by pixel coordinates, and orchestrates workflows across desktop software.

Strengths & Considerations

Core Strengths
  • Sub-pixel mouse click coordinate accuracy
  • High OSWorld benchmark score (44.8%)
  • Handles complex multi-step browser workflows
Known Limitations
  • Requires sandboxed virtual machine environment for safe execution

Token & API Pricing

Input Tokens (1M)$1.00
Output Tokens (1M)$4.00
Cached Input (1M)$0.2500
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)