Skip to main content
Live
Main content

Google bakes computer use into Gemini 3.5 Flash as a native tool

DeepMind folds its standalone agent model into Flash, letting developers build agents that drive browsers, mobile apps and desktops via one API call.

Jaeden Schafer
Editor in Chief · · 4 min read
Google logo

Google DeepMind has folded computer use into Gemini 3.5 Flash as a native built-in tool, letting developers build agents that drive browsers, mobile apps and desktop software through a single Gemini API call. The capability shipped on June 24, 2026, and replaces the standalone Gemini 2.5 computer use model that DeepMind had run as a separate endpoint. Mateo Quiros, a product manager at Google DeepMind, said the integrated version delivers the company's best performance yet on agentic computer use tasks.

The shift matters because it collapses what had been a two-model pattern — a fast generalist plus a specialist screen-driver — into one Flash-tier model that handles both. Developers can now call function calling, Search grounding, Maps grounding and computer use from the same model context without juggling handoffs. Quiros framed the launch as the moment Flash becomes a single substrate for building agents.

Computer use is now a built-in tool supported in Gemini 3.5 Flash, delivering our best performance yet for agentic computer use tasks.
Mateo Quiros, Product Manager, Google DeepMind

The pitch to enterprise buyers is long-horizon work that doesn't fit neatly into an API. DeepMind cites continuous software testing and knowledge work across professional applications — the kind of tasks that require an agent to stay coherent across dozens of clicks, form fields and page loads. In its own demo, 3.5 Flash uses computer use to crawl the Gemini app and return a categorized feature list, and in another it audits its own documentation site for accessibility issues.

Key facts

  • 01Google DeepMind made computer use a native tool inside Gemini 3.5 Flash on June 24, 2026, retiring its standalone Gemini 2.5 computer use model.
  • 02Agents built on 3.5 Flash can see, reason and act across browser, mobile and desktop environments through a single Gemini API call.
  • 03Two optional enterprise safeguards ship alongside it: human confirmation for irreversible actions and automatic task halts on detected prompt injection.
  • 04Developers can test the tool in a hosted demo environment run by Browserbase before wiring it into the Gemini Enterprise Agent Platform.

Distribution runs through two channels. Individual developers can hit it through the Gemini API; enterprise customers get it via the Gemini Enterprise Agent Platform, the managed surface Google has been building out for production agent deployments. A hosted demo environment from Browserbase lets teams kick the tires on a sandboxed browser before standing up their own infrastructure.

Security is where this kind of feature usually breaks. An agent with the keys to a live browser is one prompt injection away from clicking the wrong button, exfiltrating session cookies or wiring money to the wrong account. DeepMind says it used targeted adversarial training during model development to harden 3.5 Flash against the most common injection patterns operators see in the wild.

On top of the model-level work, Google is shipping two optional enterprise safeguard systems. One requires explicit user confirmation before the agent takes sensitive or irreversible actions — a checkout, a delete, a money movement. The other automatically halts the task if the system detects an indirect prompt injection in the page content the agent is reading. Quiros described the overall posture as defense in depth, and pushed developers to combine the model controls with sandboxing, human review and tight access scoping.

To mitigate some of the prompt injection risks for agents operating in live environments, we use targeted adversarial training for computer use in Gemini 3.5 Flash.
Mateo Quiros, Product Manager, Google DeepMind

The competitive backdrop is crowded. Anthropic shipped its own computer use capability on Claude in late 2024, and OpenAI has been pushing Operator and a sequence of agent-flavored products through 2025 and 2026. Microsoft's Copilot stack and a wave of agent-platform startups — Browserbase among them — have made browser-driving agents the most contested frontier in applied AI. Google's move is to make the capability cheap, fast and native to a model tier that developers already reach for by default.

There are real questions left open. DeepMind did not publish benchmark numbers against the prior Gemini 2.5 computer use model or against rival systems, so the claim of best-yet performance lives without a scoreboard for now. Pricing details for high-volume computer use traffic on the Enterprise Agent Platform are also light. And adversarial training narrows the prompt injection attack surface without closing it — the safeguard systems exist precisely because the model alone isn't enough.

Related · from this week
Google DeepMind's WeatherNext 3 delivers hourly forecasts at 5-kilometer resolution
Jaeden Schafer · 5 min read →

Putting computer use inside the cheapest, fastest tier of the lineup is the strategically interesting bet. Agents that have to click through real software for hours don't pencil out at premium per-token pricing, and the labs that win the agent layer will be the ones whose unit economics survive contact with long-running workflows. Google is wagering that Flash plus a native screen-driving tool is the right shape for that market — and that enterprises will buy the safeguards alongside, rather than rolling their own.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Google logo
Models

Google DeepMind's WeatherNext 3 delivers hourly forecasts at 5-kilometer resolution

The new model trains on live satellite data instead of physics simulations, cutting the six-hour lag that plagues traditional forecasting.

Jaeden Schafer5 min read
Google logo
Models

DeepMind's SIMA 2 agent plays, reasons and learns in 3D game worlds

Google DeepMind extends 15 years of games research from Atari to EVE Online with a new agent that collaborates with human players.

Jaeden Schafer4 min read
Google logo
Models

Google ships Nano Banana 2 Lite and Gemini Omni Flash to developers

DeepMind's fastest image model generates in 4 seconds at $0.034 per 1K images; Omni Flash matches Veo 3.1 Fast at $0.10 per second of video.

Jaeden Schafer5 min read