Skip to main content
Live
Main content

Google ships three new Gemini models but delays Pro update

Gemini 3.6 Flash cuts token usage by up to 17%, but the flagship Pro refresh is still stuck as OpenAI and Anthropic pull ahead.

Jaeden Schafer
Editor in Chief · · 5 min read
Google logo

Google DeepMind released three new Gemini models on Tuesday — Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber — with the headline 3.6 Flash model cutting token usage by up to 17% versus its predecessor. The releases are aimed squarely at developers running high-volume AI agents, where each token saved multiplies across billions of calls. Missing from the lineup: any update to Gemini Pro, Google's flagship, which has not been refreshed since February.

Gemini 3.6 Flash is positioned as the workhorse tier, with improved coding, knowledge, and multimodal performance at a lower per-token price than 3.5 Flash. Gemini 3.5 Flash-Lite is the cheapest option in the family, targeting cost-sensitive production workloads. Gemini 3.5 Flash Cyber is a fine-tuned variant for finding and fixing security vulnerabilities, available only to governments and trusted partners under a limited pilot.

The gap at the top of the stack is the story. Gemini Pro's last update landed in February, and in the roughly five months since, OpenAI has shipped GPT-5.5 and started rolling out GPT-5.6, while Anthropic has launched Claude Opus 4.8, Claude Sonnet 5, and expanded access to its frontier Fable 5 model. Google's flagship line has effectively sat out two release cycles from its two closest rivals.

Key facts

  • 01Google DeepMind released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber on Tuesday.
  • 02Gemini 3.6 Flash cuts token usage by up to 17% versus 3.5 Flash, lowering per-call cost.
  • 03Gemini Pro has not been updated since February; the 3.5 Pro launch is still not shipped.
  • 04OpenAI shipped GPT-5.5 and began rolling out GPT-5.6 in the same window.
  • 05Anthropic launched Claude Opus 4.8, Claude Sonnet 5, and expanded access to Fable 5.

Google itself set the expectation for a faster cadence. When 3.5 Flash launched in May, the company said the Pro version was already being used internally and would roll out the following month. That didn't happen. Bloomberg reported last week that Google is facing internal delays on 3.5 Pro after struggling to meet its own performance targets.

already being used internally, and we look forward to rolling it out next month.
Google, Company statement, May 2026

Logan Kilpatrick, product lead at Google DeepMind, said Tuesday that Gemini 3.5 Pro is currently in partner testing and the team hopes it will land soon. He also noted that DeepMind has begun what he described as its most ambitious pre-training run yet for Gemini 4 — a signal that the company may be routing more compute to the next generation rather than squeezing out a Pro release that already looks stale.

The Flash-tier releases still matter for Google's cloud business. Flash models are what most production customers actually deploy: they prioritize latency and cost over raw reasoning ceiling, and they run the agentic workloads that enterprises are scaling this year. Google framed the focus as delivering "efficiency, latency, and reliability to customers that are building AI agents at scale."

to deliver efficiency, latency, and reliability to customers that are building AI agents at scale
Google DeepMind, Official statement

That framing lines up with the pricing math. A 17% reduction in tokens on a model that already sits below Pro on cost is meaningful for any customer running Gemini in a loop — code assistants, customer support agents, document pipelines. At agent scale, single-digit percentage cuts translate into six- and seven-figure annual savings on inference bills.

The Cyber variant is the more unusual move. Fine-tuning a Flash-class model for vulnerability discovery and gating it to governments and trusted partners puts Google into the same specialized-model territory Anthropic and OpenAI have staked out with defense and intelligence customers. It's a small pilot, but it signals that the Gemini roadmap now includes vertical, access-restricted SKUs alongside the general-purpose lineup.

Related · from this week
Gemini's Spark, Daily Brief, and chat sprawl expose an AI branding problem
Jaeden Schafer · 4 min read →

The counterweight is straightforward: shipping three Flash models does not close a five-month gap at the top of the stack. Customers evaluating frontier reasoning — complex coding tasks, long-context analysis, agent orchestration that requires the smartest available model — currently have GPT-5.6 and Claude Opus 4.8 as the reference points. Every additional week that Gemini 3.5 Pro sits in partner testing is a week competitors get to define what a 2026 flagship looks like.

Google's bet appears to be that the volume tier is where the revenue actually lives, and that leapfrogging directly to Gemini 4 is worth more than a hurried Pro release that would ship behind the rivals anyway. That is a defensible strategy if Gemini 4 arrives on a competitive timeline. It is a much harder story to tell customers if the Pro gap stretches into the fall and Gemini 4 slips. Flash cuts costs today; only a credible flagship keeps developers from routing their most valuable workloads to someone else.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Google logo
Analysis

Gemini's Spark, Daily Brief, and chat sprawl expose an AI branding problem

Google's Gemini app now houses three separately branded features — a pattern Anthropic and OpenAI repeat, and Apple deliberately avoids.

Jaeden Schafer4 min read
Google logo
Careers

Google loses Gemini architects Adler and Pritzel to Anthropic

Two more Gemini researchers join a wave of defections to Anthropic and OpenAI, weeks after Google paid $2.7B to bring Noam Shazeer back.

Jaeden Schafer4 min read
Google logo
Business

Google cuts AI Plus to $4.99 a month, opening a US price war

Google halves storage cost and undercuts ChatGPT Plus by 75%, importing a pricing fight that started in India.

Jaeden Schafer5 min read