Introducing Gemini Robotics On-DeviceGoogleReleased July 1, 2025

Gemini Robotics On-Device

On-device vision-language-action foundation model for real-time robotic perception and physical manipulation.

multimodal-realtimeProprietary APIOn-Device Edge EmbeddedContext: 16.384K (16,384 tokens)

Technical Specifications

Architecture Type
Edge Vision-Language-Action (VLA) Model
Total Parameters
Compact Edge VLA
Context Window
16.384K (16,384 tokens)
Max Output Tokens
2.048K (2,048 tokens)
Knowledge Cutoff
March 2025
Supported Modalities
vision, text, actions
License & Access
Google Robotics Research License

Benchmark Evaluations

Latency Ms
45
Manipulation Success Rate
89.2

Deep Architectural Overview

Gemini Robotics On-Device enables autonomous physical robots to perceive environments, reason about spatial constraints, and output low-latency trajectory control signals directly on edge hardware.

Strengths & Considerations

Core Strengths
  • Real-time 45ms control frequency
  • Zero-shot generalization to unseen objects
  • Full on-device privacy and reliability without cloud connection
Known Limitations
  • Restricted to specialized robotic actuators and edge TPUs

Token & API Pricing

Input Tokens (1M)Contact
Output Tokens (1M)Contact
Pricing is verified directly against Google's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Google and comparable reasoning engines.

Browse all models
Google

The latest evolution in the Gemini 3 family, delivering state-of-the-art software engineering (73.7% DeepSWE) and agentic enterprise knowledge workflows.

1.049M ctx$0.75 / 1M tok ($1.50 reg)
Google

Universal omnimodal generation model that accepts any combination of text, audio, image, and video to generate any combination of outputs.

1.049M ctx$0.60 / 1M tok (Universal Omni)