MODEL
Gemini 3.5 Flash
Overview
Gemini 3.5 Flash is Google’s I/O 2026 launch of the Flash-tier model in the 3.5 family, positioned explicitly at long-horizon agentic workflows rather than chat. Pricing of $1.50 / $9.00 per million input/output tokens is roughly 25% below Gemini 3.1 Pro’s $2.00 / $12.00 — pulling a frontier-tier coding model into Flash pricing for agent builders who were already managing a Pro-tier cost line. Notably absent from Google’s comparison materials: any Pro 3.5 number.
Timeline
- 2026-05-20-AI-Digest — Google unveils Gemini 3.5 Flash at I/O 2026 with vendor-reported 76.2% on Terminal-Bench 2.1 (vs 70.3% for Gemini 3.1 Pro), 1656 Elo on GDPval-AA, and a beat on MCP Atlas — all on Google’s own benchmarks. Pricing is $1.50 / $9.00 per million input/output tokens vs 3.1 Pro’s $2.00 / $12.00 — about 25% cheaper on both sides, not the “half cost” some early coverage carried. Framing pulls a frontier-tier coding model into Flash pricing; the arc to watch is whether Cursor and similar IDEs swap their default Flash tier. Cursor Composer 2.5‘s $0.50 / $2.50 still undercuts Flash on input pricing, so the practitioner heuristic “under $1 per agentic task” hasn’t moved — it’s just been joined by a new frontier-lab option at a similar order of magnitude.
- 2026-06-25-AI-Digest — Google extends its computer-use agent API to Gemini 3.5 Flash (Google blog post hits HN at 194 pts / 124 cmts). The narrow read: computer-use moves from the Pro tier down to the cheaper/faster Flash tier. The structural read the digest carries: this makes browser-and-OS-controlling agents economically viable for high-volume workloads, intensifying the agent-platform race with Anthropic and OpenAI — the cost-per-task floor on browser-controlling agents drops with the Flash tier rather than waiting for a Pro-tier price cut.
- 2026-06-26-AI-Digest — Google folds Computer Use directly into Gemini 3.5 Flash as a native capability, replacing the previous Gemini-2.5-Computer-Use-Preview spinoff model. On the OSWorld desktop-agent benchmark Gemini 3.5 Flash with Computer Use lands at 78.4, between Claude Opus 4.8 (83.4) and GPT-5.4 mini (72.1). Narrow read: Computer Use is no longer a separate-model side bet — it’s a capability of the cheap-tier flagship, which collapses the “do I pay for the dedicated agent model” decision. Structural read worth carrying: for high-volume browser-and-desktop agent workloads, Gemini 3.5 Flash is now the price-performance default until Anthropic drops a Haiku-tier computer-use SKU or OpenAI inverts the gap with a cheaper-tier release. Pairs with the Claude Tag launch from 2026-06-23-AI-Digest — the agent-platform race is layered (“agent identity in collaboration surface” vs “agent that can drive your desktop on price”) and the layers are not directly substitutable.
- Computer Use Folded Into the Flash Tier as a Native Capability (June 26, 2026): The Gemini-2.5-Computer-Use-Preview spinoff model retires; Computer Use becomes a capability of the cheap-tier flagship at OSWorld 78.4 — between Claude Opus 4.8 at 83.4 and GPT-5.4 mini at 72.1. The Flash tier is now the price-performance default for high-volume browser-and-desktop agent workloads. Whether Anthropic follows with a Haiku-tier computer-use SKU, or OpenAI inverts the gap with a cheaper-tier release, is the watch item for the next quarter.
- 2026-07-22-AI-Digest — Google / DeepMind ship three Flash-tier Gemini models today — Gemini 3.6 Flash (live in Vertex Model Garden, cutting reported token usage by up to 17%), Gemini 3.5 Flash-Lite (smallest tier), and Gemini 3.5 Flash Cyber (security-tuned to find, validate, and patch vulnerabilities) — with no accompanying Gemini 3.5 Pro. Aliases the Gemini 3.6 Flash release to this note (
[[Gemini 3.5 Flash|Gemini 3.6 Flash]]) rather than spawning a new topic note — the 3.6 Flash refresh is a token-efficiency drop on the same family shape rather than a distinct capability tier. Bloomberg’s earlier “held back for coding-benchmark targets” reporting on the Pro delay fits today’s shipment shape: Google shipped the tier where the bar was met and skipped the tier where it wasn’t. Aider polyglot top-5 still hasgemini-2.5-pro-preview-06-05at #4 — a preview Pro line rather than a shipped 3.5 Pro. Structural read the corpus carries: a Flash refresh with no accompanying Pro tier is not routine cadence — it is the “held back for coding targets” story showing up in the release schedule, and Flash Cyber is a sideways move into the security-tuned lane Anthropic (Claude Code Security) and OpenAI (GPT-5.5 Cyber previously) already sit in. 90-day watch: whether Gemini 3.5 Pro lands before end of Q3.
- Flash 3.6 Refresh + Flash-Lite + Flash Cyber Ship With No 3.5 Pro (July 22, 2026): Three Flash-tier ships in a single release window — Gemini 3.6 Flash (up to 17% token-usage cut on Vertex Model Garden), Gemini 3.5 Flash-Lite (smallest tier), Gemini 3.5 Flash Cyber (security-tuned). Notably absent: any Gemini 3.5 Pro. Aider polyglot top-5’s #4 slot still shows
gemini-2.5-pro-preview-06-05— a preview line, not a shipped Pro tier. The Bloomberg “held back for coding targets” framing is now the base case; Flash Cyber is Google’s entry in the security-tuned-model race Anthropic (Claude Code Security) and OpenAI (GPT-5.5 Cyber) already sit in. 3.6 Flash / Flash-Lite / Flash Cyber alias to this family note rather than each getting their own topic note under §5b’s conservative threshold.
- 2026-07-23-AI-Digest — DeepMind Gemini 3.5 Flash Cyber shipped 2026-07-21 as a gated pilot for governments and trusted partners — tuned to find, validate, and patch vulnerabilities, integrated with the CodeMender agent (per Help Net Security coverage of the DeepMind post). The Flash Cyber launch pairs with Cisco‘s same-slot Antares 350M/1B Apache-2.0 open-weight cybersec release on Hugging Face as the two shapes of the bifurcating security-lane market: open-weight cost-optimised (Cisco Antares) for practitioner and enterprise adoption, and sovereign-gated capability-maximum (Gemini 3.5 Flash Cyber) for state and critical-infrastructure buyers. Narrow read: extends the 2026-07-22-AI-Digest Flash Cyber shipment thread with the practitioner-press write-ups that surface the gated-pilot distribution shape. Structural read the corpus carries: the vulnerability-detection task is splitting into two market shapes distinct from the general-purpose-frontier lane, and belongs on the agent-security MOC’s radar as its own thread rather than as a footnote on the frontier-model story. No fresh Flash refresh today beyond yesterday’s cadence; log as bifurcation-thread extension with a fresh distribution-shape datapoint.
- 2026-07-25-AI-Digest — Gemini 3.5 Flash Cyber recovered as a defensive-AI enterprise/gov sales motion positioning — the July 21 limited-pilot release (2026-07-22-AI-Digest shipment, 2026-07-23-AI-Digest practitioner-press pickup) is now framed by the corpus as fitting alongside Anthropic‘s Alberta cybersecurity case study earlier this month as the vendor-side beginnings of a defensive-AI enterprise/gov sales motion. The pitch is “your defenders can move at model speed, too” — a direct answer to the offensive-AI narrative that the July 22 GPT-5.6 Sol / Hugging Face ExploitGym incident (post-mortem in 2026-07-24-AI-Digest) crystallised into a real market anxiety. Narrow read: this is a distribution move (gated-pilot access surface + CodeMender-integrated fine-tune wrapper), not a capabilities move — the Flash-tier base model is unchanged. Structural read the corpus carries: expect the same defensive-AI sales-motion play from Anthropic and OpenAI within 30–60 days.
- 2026-07-27-AI-Digest — Gemini 3.5 Flash Cyber gets its cleanest single-line corpus framing today — “cyber defense as the next contested small-model vertical, in the wake of the Hugging Face / OpenAI incident and Anthropic‘s Alberta-government cybersecurity work.” The Jul 21 ship (lightweight 3.5 Flash variant fine-tuned to find, validate, and patch vulnerabilities, released alongside 3.6 Flash and 3.5 Flash-Lite, integrated with CodeMender in a gated pilot) now anchors the corpus’s read of small, specialised cyber-defence models as a distinct product category alongside general reasoning models. No fresh 3.5 Flash cadence today; the news beat is the corpus-level framing hardening around the earlier ship rather than a new release.
- 2026-08-11-AI-Digest — Gemini 3.5 Flash Cyber lands as Google‘s Aug leg of the three-lab cyber triopoly that crystallises with OpenAI‘s Aug 10 GPT-5.6-Cyber launch under Daybreak Red and Anthropic‘s Claude Mythos 5 under Project Glasswing — three of the four US frontier labs now shipping purpose-built cyber SKUs to gated enterprise defenders within a four-month window. Corpus framing to carry: each vendor’s Aug cyber SKU has different names and positioning (Claude Mythos 5 is Anthropic’s; Gemini 3.5 Flash Cyber is Google’s gov-and-trusted-partner variant under the AI Threat Defense umbrella); the “OpenAI joins the club” framing flattens meaningful design differences. No fresh 3.5 Flash Cyber cadence today; the news beat is the triopoly-framing anchor around the earlier ship — Meta remains the outlier and whether it ships a cyber-tuned Llama variant is the four-lab-vs-three-lab question for the next quarter.
- Third Leg of the Three-Lab Cyber Triopoly Alongside GPT-5.6-Cyber and Claude Mythos 5 Inside a Four-Month Window (August 11, 2026): The Jul 21 Gemini 3.5 Flash Cyber ship now reads as the earliest of the three US frontier-lab cyber SKUs that crystallise into a real triopoly with GPT-5.6-Cyber under Daybreak Red on Aug 10 — Claude Mythos 5 under Project Glasswing is the third leg. Load-bearing corpus framing: the “OpenAI joins Anthropic” framing is one lab behind — the accurate frame is three labs, three purpose-built cyber SKUs, three different names + gating shapes (Google under AI Threat Defense for gov-and-trusted-partners, Anthropic under Project Glasswing, OpenAI under Daybreak Blue/Red). Meta the outlier; four-lab-vs-three-lab is the next-quarter question.
Key Developments
-
Flash-Tier Reframed for Agentic Workloads: Google’s explicit positioning shifts the Flash tier from “cost-optimized chat” to “long-horizon agent runs.” If the Terminal-Bench delta over 3.1 Pro survives independent evaluation, this is the first major frontier-lab Flash-tier model marketed primarily on coding-agent throughput rather than per-token cost.
-
Treat the Benchmarks as Google-Controlled: Every datapoint at launch is on Google’s own evaluation harness with Google’s own selection of comparators (no Pro 3.5 number, no third-party replication yet). The honest read holds the numbers as directional until independent benchmarks land.
-
Pricing Anchors but Does Not Reset the Floor: Cursor Composer 2.5 at $0.50 / $2.50 remains the cheaper agentic option on input pricing. The frontier-lab Flash tier is now in the same order of magnitude as IDE-vendor in-house models — a real convergence, but not a price floor break.