- Google released three models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber (paired with the CodeMender security agent).
- 3.6 Flash uses 17% fewer output tokens than 3.5 Flash per the Artificial Analysis Index — up to 65% fewer on DeepSWE — at a lower price of $1.50/1M input and $7.50/1M output tokens.
- Benchmarks improved across the board: DeepSWE 49% vs 37%, MLE Bench 63.9% vs 49.7%, OSWorld-Verified 83.0% vs 78.4%.
- Gemini 3.5 Pro is still only testing with partners, and Google says pre-training has started on Gemini 4.
What Happened
Google introduced Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, positioning the Flash series at the efficiency-quality sweet spot for running AI agents at scale, according to the announcement published July 21, 2026 by Tulsee Doshi, Senior Director of Product Management, on behalf of the Gemini team.
Why It Matters
The pitch is cost per agentic task, not raw capability. Production agents burn tokens across many reasoning steps and tool calls, so a model that is both cheaper per token and less verbose compounds savings. Notably, the frontier model developers have been waiting for — Gemini 3.5 Pro — is still only testing with partners, with broad availability promised “as soon as it’s ready,” while Google says it has already begun its most ambitious pre-training run yet for Gemini 4.
Technical Details
Gemini 3.6 Flash consumes 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65% fewer on some benchmarks such as DeepSWE by Datacurve, while taking fewer reasoning steps and tool calls on multi-step workflows. Pricing is $1.50 per million input tokens and $7.50 per million output tokens. Performance gains over 3.5 Flash include DeepSWE (49% vs 37%), MLE Bench (63.9% vs 49.7%), OSWorld-Verified (83.0% vs 78.4%), and GDPval-AA v2 (1421 vs 1349); computer use is now a built-in client-side tool via the Gemini API and Gemini Enterprise. 3.5 Flash-Lite delivers 350 output tokens per second as the fastest, most cost-effective 3.5-class model, and 3.5 Flash Cyber pairs a specialized cyber-focused model with the CodeMender code-security agent. 3.6 Flash ships with Frontier Safety safeguards for chemical, biological, radiological, and nuclear (CBRN) and cyber-offense misuse, which Google says make it substantially more jailbreak-resistant while minimizing refusals for beneficial uses.
Who’s Affected
The release targets developers and enterprises running production agents, where token efficiency directly sets operating cost. Google cites customers including Hebbia and Harvey finding 3.6 Flash particularly capable at multimodal work such as document parsing, chart and data analysis, and report drafting. The cyber-specialized variant aims at security teams that need frontier-competitive performance in a narrower, cheaper package.
What’s Next
The stated roadmap is Gemini 3.5 Pro reaching broad availability once ready, and Gemini 4 in pre-training. The benchmark figures come from Google and the Artificial Analysis Index rather than independent replication, so the practical test is whether the claimed token savings hold on real agentic workloads — which is exactly where the cost argument either pays off or doesn’t.