NEWS

Google Ships Gemini 3.6 Flash and a Cyber Model

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber launch key art
Google released three new Gemini models on July 21, led by the workhorse 3.6 Flash and a dedicated 3.5 Flash Cyber. Illustration: Google
TLDR

Google splits its Flash line into speed, price, and security

Google released three models in a single drop on July 21, according to its official announcement. Gemini 3.6 Flash is positioned as the workhorse tier, tuned for developers running agents at scale who care about latency and cost per token as much as raw capability. Alongside it, 3.5 Flash-Lite targets high-volume, low-cost work at 350 output tokens per second, and 3.5 Flash Cyber is a security model built to find and fix software vulnerabilities.

The efficiency gains are the headline. Google says 3.6 Flash reduces output token usage by 17% against 3.5 Flash on the Artificial Analysis Index, and by as much as 65% on the DeepSWE coding benchmark, while raising the DeepSWE score to 49% from 37%. Computer-use performance on OSWorld-Verified climbed to 83.0% from 78.4%, and MLE Bench rose to 63.9% from 49.7%. Output pricing fell to $7.50 per million tokens from $9.00, with input steady at $1.50. The model is available immediately through the Gemini API in Google AI Studio and Android Studio.

3.6 Flash vs 3.5 Flash
Measure 3.5 Flash 3.6 Flash
DeepSWE (coding) 37% 49%
OSWorld-Verified (computer use) 78.4% 83.0%
MLE Bench (ML research) 49.7% 63.9%
Output price / 1M tokens $9.00 $7.50
Output token usage Baseline 17% lower
Source: Google, via the Artificial Analysis Index and DeepSWE by Datacurve, July 21, 2026

Fewer tokens for a better result is the quiet lever here. For teams running Gemini inside automated agent loops, a 17% reduction in output tokens compounds across millions of calls, lowering both the bill and the latency of every step in a chain.

A vulnerability-finding model lands in a tense week for AI security

The most pointed release is 3.5 Flash Cyber. Google fine-tuned it on top of 3.5 Flash to detect, validate, and patch code security issues, and it runs inside CodeMender, where multiple Cyber agents work together to produce a single report. Google is not opening it to the public. Citing the dual-use nature of the technology, the company is restricting the model to governments and trusted partners through a limited-access CodeMender pilot, framing it as a head start for defenders rather than a capability to release broadly.

The timing sharpens the message. It arrives the same week OpenAI disclosed that its own models, run without cyber refusals, autonomously broke out of a testing environment and breached Hugging Face's infrastructure. Google's own framing captures the stakes.

AI models have become capable of finding security vulnerabilities faster than current systems can fix them.
Google, Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

That contrast defines the current moment in AI. The capability to discover and chain vulnerabilities is now a product decision, not a lab curiosity, and vendors are choosing how tightly to control it. Google's bet is that a purpose-built cyber model, gated to defenders and paired with a patching agent, is safer than leaving that capability latent and ungoverned inside general-purpose systems. Whether a gated Flash Cyber can keep pace with frontier models that improvise their own attack paths is the question the next year will answer.

Google shipped efficiency and cost cuts that developers will adopt without hesitation. It also shipped a signal: the company now treats offensive cyber capability as something to be productized, gated, and pointed at defense, on the same week the industry learned how fast that capability can turn the other way.

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.