Google released three new Gemini models today aimed squarely at the economics of running AI agents in production. Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber-tuned 3.5 Flash Cyber variant all target the same problem: agentic workflows that burn through tokens faster than they generate value. The headline number is a 17% reduction in output token usage on 3.6 Flash versus 3.5 Flash, measured on the Artificial Analysis Index, and up to 65% on specific benchmarks like DeepSWE by Datacurve.
The pricing moves in the same direction. Gemini 3.6 Flash lists at $1.50 per million input tokens and $7.50 per million output tokens, below its predecessor, while 3.5 Flash-Lite drops to $0.30 and $2.50 per million respectively. Flash-Lite pushes throughput to 350 output tokens per second by Artificial Analysis's measurement, which Google is pitching as the fastest model in the 3.5 series for high-volume agentic search and document processing.
“Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance.”— Tulsee Doshi, Senior Director, Product Management, Gemini
Tulsee Doshi, senior director of product management for Gemini, framed the release as a response to production feedback rather than a capability jump. The Flash line has always been Google's efficiency tier below Pro, but this generation leans harder into the agent use case, where a single task can spawn dozens of tool calls and reasoning steps.
Key facts
- 01Gemini 3.6 Flash consumes 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, with reductions up to 65% on DeepSWE.
- 023.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens, below 3.5 Flash.
- 033.5 Flash-Lite runs at 350 output tokens per second and prices at $0.30/$2.50 per million input/output tokens.
- 043.6 Flash scored 83.0% on OSWorld-Verified vs. 78.4% for 3.5 Flash, and 49% on DeepSWE vs. 37%.
- 05Google has started pre-training Gemini 4 and is testing Gemini 3.5 Pro with partners ahead of a broader release.
On benchmarks, 3.6 Flash hit 49% on DeepSWE versus 37% for 3.5 Flash, 63.9% on MLE Bench versus 49.7%, and 83.0% on OSWorld-Verified versus 78.4%. It also scored 1421 on GDPval-AA v2 against 3.5 Flash's 1349. Computer use is now a built-in client-side tool via the Gemini API and Gemini Enterprise, which matters for the OSWorld-style tasks where the model has to click, type, and navigate real applications.
Flash-Lite's gains against the prior generation are steeper. It scored 54% on Terminal-Bench 2.1 against 31% for 3.1 Flash-Lite, 72.2% on GDM-MRCR v2 against 60.1%, and 1140 on GDPval-AA v2 against 642. Google is claiming it beats the previous full-size 3 Flash on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), which would let developers downshift to the Lite tier for workloads that used to require Flash. Hebbia and Harvey were named as customers finding 3.6 Flash particularly useful for document parsing and chart analysis.
The third model, 3.5 Flash Cyber, is a more targeted release. It is a fine-tuned variant built on 3.5 Flash for finding and fixing cybersecurity vulnerabilities, and it will ship exclusively inside CodeMender, Google's code security agent. The system uses multiple 3.5 Flash Cyber agents in orchestration to produce a single combined vulnerability report, and Google says it reaches competitive frontier performance on the CyberGym benchmark.
“AI models have become capable of finding security vulnerabilities faster than current systems can fix them.”— Tulsee Doshi, Senior Director, Product Management, Gemini
Google is not making Flash Cyber generally available. The model will be limited to governments and trusted partners via CodeMender under a pilot program, citing the dual-use risk of a model trained specifically to find security holes. That is a notably tighter distribution posture than Google's other Flash releases, which are landing in Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, the Gemini app, and Google Search from today.
The 3.6 Flash safety story is built around what Google calls Frontier Safety safeguards in CBRN and cyber offense domains, with training aimed at resisting jailbreaks while minimizing refusals on benign requests. That balance has been a persistent complaint against Gemini's earlier releases, where over-refusal on legitimate developer queries hurt adoption.
The obvious missing piece is Gemini 3.5 Pro, which Google says is currently in partner testing and will ship when ready. Meanwhile, the company confirmed it has started what Doshi called its most ambitious pre-training run yet for Gemini 4. That framing puts the Flash refresh in context: incremental efficiency work on the current generation while the next flagship is still cooking.
Skeptics will note that Google is grading its own homework here. The Artificial Analysis Index is a third-party benchmark, but many of the comparisons the company highlights are Gemini-versus-Gemini rather than against Claude, GPT-5, or the open-weight models from Moonshot and Alibaba that we covered last week. There is no head-to-head on SWE-Bench or OSWorld against the current frontier from Anthropic or OpenAI in Google's post, which leaves the competitive picture partially unanswered.
The strategic read is that Google is fighting the agent-economics war rather than the capability war. When an autonomous coding agent chews through millions of tokens on a single ticket, a 17% output reduction combined with a price cut compounds into real margin for the customer and real volume for Google. Flash-Lite beating the older full-Flash on agentic evals lets Google pull developers down the pricing ladder without losing the workload, which is how you monetize an inference layer at scale. The Gemini 4 comment is the reminder that this is the interim, not the destination.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




