Gemini 3.6 Flash is here — and 3.5 Pro is still waiting

Google shipped a sharper Flash workhorse, a faster Lite tier, and 3.5 Flash Cyber inside CodeMender. Refresh Gemini web today; update the mobile apps if the new models are missing — and stop holding your breath for 3.5 Pro just yet.

Read time: ~7 minutes

Gemini 3.6 Flash and 3.5 Flash-Lite are live for everyday use. Google's July 21 announcement also introduces 3.5 Flash Cyber inside CodeMender — a multi-agent security stack for finding and patching vulnerabilities. The consumer Flash models are the efficiency layer for production agents: better coding and knowledge work, fewer tokens, lower latency — without waiting for the Pro tier a lot of people expected next.

If you use Gemini in the browser, a refresh is usually enough to see the new models. On phone, open the App Store or Google Play, update the Gemini app to the latest build, then reopen — older installs can lag a release cycle behind web.

Try it now: Gemini web → refresh. Mobile → update Gemini from App Store / Play, then relaunch. Developers: Gemini API via Google AI Studio; enterprises: Gemini Enterprise Agent Platform.

What actually shipped

3.5 Flash Cyber inside CodeMender

CodeMender is Google's code-security agent system. Multiple 3.5 Flash Cyber agents work together to discover vulnerabilities, validate them, and write patches for complex software issues — then merge into a single combined report. Google positions the stack as competitive on CyberGym-style benches while staying efficient enough to run at scale.

Because dual-use risk is real, Flash Cyber + CodeMender is rolling out as a limited-access pilot for governments and trusted partners — a head start for defenders, not a public model picker option for vibe coders. Worth knowing it exists; don't expect it in your Gemini mobile app update.

Google also notes Gemini 3.5 Pro is still testing with partners, with broader availability "as soon as it's ready," while the lab has already started its most ambitious pre-training run yet for Gemini 4.

Where 3.5 Pro went (and why the wait is rational)

Plenty of builders were waiting for Gemini 3.5 Pro as the next headline drop. Instead Google pushed Flash efficiency first. That reads less like "Pro is cancelled" and more like "Pro isn't ready to sit at the top of the table."

Look at the last few months alone: Opus 4.8, Fable 5, GPT-5.6, Kimi K3, and Cursor Grok 4.5 all landed in the same window. Google cannot casually ship a mid-tier Pro into that pile and call it a win. The AI race is on — and as users we benefit: better outputs now, and real pressure toward lower cost over time.

Pricing: where Flash sits

Flash is not trying to beat Opus on price-per-flagship-token theater. It's the high-volume agent layer — cheaper than prior Flash, and in the same neighborhood as budget coding models when you care about loops, not single-shot brilliance.

Model Input / 1M Output / 1M Notes
Gemini 3.6 Flash $1.50 $7.50 Workhorse; fewer tokens than 3.5 Flash
Gemini 3.5 Flash-Lite $0.30 $2.50 ~350 tok/s (Artificial Analysis)
Grok 4.5$2$6Cursor / SpaceXAI API
GPT-5.6 Luna$1$6Budget OpenAI tier
Claude Opus 4.8$5$25Prior-gen flagship

Benchmarks: honest placement, not a clean sweep

Google published gains for 3.6 Flash vs its own 3.5 Flash (DeepSWE, OSWorld-Verified, MLE Bench, GDPval) and strong Flash-Lite gains vs older Lite / Flash baselines. Cross-vendor tables are messier — different harnesses, thinking levels, and "max" settings. Below we mix Google's cited scores with numbers we already published for Grok, Opus, and Fable. Treat dashes as "not in that source," not as zeros.

Benchmark 3.6 Flash 3.5 Flash Flash-Lite Grok 4.5 Opus 4.8 Fable 5
DeepSWE 49%vs 3.5 Flash 37% 62%high 55.8%max 66.1%max
OSWorld-Verified 83.0% 78.4% 74.0%
Terminal-Bench 2.1 54%vs 3.1 Lite 83.3% 78.9% 84.3%
SWE-Bench Pro 54.2%vs Gemini 3 Flash 64.7%high 69.2%max 80.3%
MLE Bench 63.9% 49.7%

Read: 3.6 Flash is a clear step up inside Google's Flash line — especially token efficiency and computer-use / multimodal knowledge work. It is not automatically the new SWE-Bench Pro king; Fable 5 and Opus still own several coding peaks we track, and Grok 4.5 still wins on the "near-frontier + cheap loops" story for Cursor users. Flash-Lite is the scale dial: slower brains, absurd throughput and price for agent swarms.

How this fits the recent model posts

Kimi K3 showed up in our last newsletter digest — open-weights noise in the same race, no standalone post yet.

Who should use what

Bottom line

Google shipped the Flash upgrades users can run today, held Pro until it can compete at the top, and already started the Gemini 4 pre-train. Refresh Gemini web, update the mobile app, try 3.6 Flash on a real agent loop, and keep following the race — cheaper, sharper models are the consumer upside of everyone fighting for the leaderboard.

Sources: Google Gemini blog (Jul 21, 2026); cross-vendor benches from our prior Grok / GPT-5.6 coverage. Harnesses differ — don't treat the grid as a single official league table.

Disclosure: Some links above are referral or partner links (marked on our Tools page).

Questions? Get in touch — or subscribe for the next AI news post.

← All posts