MODEL

Gemini 3.5 Flash

modeltopic-notegoogle

Overview

Gemini 3.5 Flash is Google’s I/O 2026 launch of the Flash-tier model in the 3.5 family, positioned explicitly at long-horizon agentic workflows rather than chat. Pricing of $1.50 / $9.00 per million input/output tokens is roughly 25% below Gemini 3.1 Pro’s $2.00 / $12.00 — pulling a frontier-tier coding model into Flash pricing for agent builders who were already managing a Pro-tier cost line. Notably absent from Google’s comparison materials: any Pro 3.5 number.

Timeline

  • 2026-05-20-AI-Digest — Google unveils Gemini 3.5 Flash at I/O 2026 with vendor-reported 76.2% on Terminal-Bench 2.1 (vs 70.3% for Gemini 3.1 Pro), 1656 Elo on GDPval-AA, and a beat on MCP Atlas — all on Google’s own benchmarks. Pricing is $1.50 / $9.00 per million input/output tokens vs 3.1 Pro’s $2.00 / $12.00 — about 25% cheaper on both sides, not the “half cost” some early coverage carried. Framing pulls a frontier-tier coding model into Flash pricing; the arc to watch is whether Cursor and similar IDEs swap their default Flash tier. Cursor Composer 2.5‘s $0.50 / $2.50 still undercuts Flash on input pricing, so the practitioner heuristic “under $1 per agentic task” hasn’t moved — it’s just been joined by a new frontier-lab option at a similar order of magnitude.
  • 2026-06-25-AI-DigestGoogle extends its computer-use agent API to Gemini 3.5 Flash (Google blog post hits HN at 194 pts / 124 cmts). The narrow read: computer-use moves from the Pro tier down to the cheaper/faster Flash tier. The structural read the digest carries: this makes browser-and-OS-controlling agents economically viable for high-volume workloads, intensifying the agent-platform race with Anthropic and OpenAI — the cost-per-task floor on browser-controlling agents drops with the Flash tier rather than waiting for a Pro-tier price cut.
  • 2026-06-26-AI-DigestGoogle folds Computer Use directly into Gemini 3.5 Flash as a native capability, replacing the previous Gemini-2.5-Computer-Use-Preview spinoff model. On the OSWorld desktop-agent benchmark Gemini 3.5 Flash with Computer Use lands at 78.4, between Claude Opus 4.8 (83.4) and GPT-5.4 mini (72.1). Narrow read: Computer Use is no longer a separate-model side bet — it’s a capability of the cheap-tier flagship, which collapses the “do I pay for the dedicated agent model” decision. Structural read worth carrying: for high-volume browser-and-desktop agent workloads, Gemini 3.5 Flash is now the price-performance default until Anthropic drops a Haiku-tier computer-use SKU or OpenAI inverts the gap with a cheaper-tier release. Pairs with the Claude Tag launch from 2026-06-23-AI-Digest — the agent-platform race is layered (“agent identity in collaboration surface” vs “agent that can drive your desktop on price”) and the layers are not directly substitutable.
  1. Computer Use Folded Into the Flash Tier as a Native Capability (June 26, 2026): The Gemini-2.5-Computer-Use-Preview spinoff model retires; Computer Use becomes a capability of the cheap-tier flagship at OSWorld 78.4 — between Claude Opus 4.8 at 83.4 and GPT-5.4 mini at 72.1. The Flash tier is now the price-performance default for high-volume browser-and-desktop agent workloads. Whether Anthropic follows with a Haiku-tier computer-use SKU, or OpenAI inverts the gap with a cheaper-tier release, is the watch item for the next quarter.
  • 2026-07-22-AI-DigestGoogle / DeepMind ship three Flash-tier Gemini models today — Gemini 3.6 Flash (live in Vertex Model Garden, cutting reported token usage by up to 17%), Gemini 3.5 Flash-Lite (smallest tier), and Gemini 3.5 Flash Cyber (security-tuned to find, validate, and patch vulnerabilities) — with no accompanying Gemini 3.5 Pro. Aliases the Gemini 3.6 Flash release to this note ([[Gemini 3.5 Flash|Gemini 3.6 Flash]]) rather than spawning a new topic note — the 3.6 Flash refresh is a token-efficiency drop on the same family shape rather than a distinct capability tier. Bloomberg’s earlier “held back for coding-benchmark targets” reporting on the Pro delay fits today’s shipment shape: Google shipped the tier where the bar was met and skipped the tier where it wasn’t. Aider polyglot top-5 still has gemini-2.5-pro-preview-06-05 at #4 — a preview Pro line rather than a shipped 3.5 Pro. Structural read the corpus carries: a Flash refresh with no accompanying Pro tier is not routine cadence — it is the “held back for coding targets” story showing up in the release schedule, and Flash Cyber is a sideways move into the security-tuned lane Anthropic (Claude Code Security) and OpenAI (GPT-5.5 Cyber previously) already sit in. 90-day watch: whether Gemini 3.5 Pro lands before end of Q3.
  1. Flash 3.6 Refresh + Flash-Lite + Flash Cyber Ship With No 3.5 Pro (July 22, 2026): Three Flash-tier ships in a single release window — Gemini 3.6 Flash (up to 17% token-usage cut on Vertex Model Garden), Gemini 3.5 Flash-Lite (smallest tier), Gemini 3.5 Flash Cyber (security-tuned). Notably absent: any Gemini 3.5 Pro. Aider polyglot top-5’s #4 slot still shows gemini-2.5-pro-preview-06-05 — a preview line, not a shipped Pro tier. The Bloomberg “held back for coding targets” framing is now the base case; Flash Cyber is Google’s entry in the security-tuned-model race Anthropic (Claude Code Security) and OpenAI (GPT-5.5 Cyber) already sit in. 3.6 Flash / Flash-Lite / Flash Cyber alias to this family note rather than each getting their own topic note under §5b’s conservative threshold.
  • 2026-07-23-AI-DigestDeepMind Gemini 3.5 Flash Cyber shipped 2026-07-21 as a gated pilot for governments and trusted partners — tuned to find, validate, and patch vulnerabilities, integrated with the CodeMender agent (per Help Net Security coverage of the DeepMind post). The Flash Cyber launch pairs with Cisco‘s same-slot Antares 350M/1B Apache-2.0 open-weight cybersec release on Hugging Face as the two shapes of the bifurcating security-lane market: open-weight cost-optimised (Cisco Antares) for practitioner and enterprise adoption, and sovereign-gated capability-maximum (Gemini 3.5 Flash Cyber) for state and critical-infrastructure buyers. Narrow read: extends the 2026-07-22-AI-Digest Flash Cyber shipment thread with the practitioner-press write-ups that surface the gated-pilot distribution shape. Structural read the corpus carries: the vulnerability-detection task is splitting into two market shapes distinct from the general-purpose-frontier lane, and belongs on the agent-security MOC’s radar as its own thread rather than as a footnote on the frontier-model story. No fresh Flash refresh today beyond yesterday’s cadence; log as bifurcation-thread extension with a fresh distribution-shape datapoint.
  • 2026-07-25-AI-DigestGemini 3.5 Flash Cyber recovered as a defensive-AI enterprise/gov sales motion positioning — the July 21 limited-pilot release (2026-07-22-AI-Digest shipment, 2026-07-23-AI-Digest practitioner-press pickup) is now framed by the corpus as fitting alongside Anthropic‘s Alberta cybersecurity case study earlier this month as the vendor-side beginnings of a defensive-AI enterprise/gov sales motion. The pitch is “your defenders can move at model speed, too” — a direct answer to the offensive-AI narrative that the July 22 GPT-5.6 Sol / Hugging Face ExploitGym incident (post-mortem in 2026-07-24-AI-Digest) crystallised into a real market anxiety. Narrow read: this is a distribution move (gated-pilot access surface + CodeMender-integrated fine-tune wrapper), not a capabilities move — the Flash-tier base model is unchanged. Structural read the corpus carries: expect the same defensive-AI sales-motion play from Anthropic and OpenAI within 30–60 days.
  • 2026-07-27-AI-DigestGemini 3.5 Flash Cyber gets its cleanest single-line corpus framing today — “cyber defense as the next contested small-model vertical, in the wake of the Hugging Face / OpenAI incident and Anthropic‘s Alberta-government cybersecurity work.” The Jul 21 ship (lightweight 3.5 Flash variant fine-tuned to find, validate, and patch vulnerabilities, released alongside 3.6 Flash and 3.5 Flash-Lite, integrated with CodeMender in a gated pilot) now anchors the corpus’s read of small, specialised cyber-defence models as a distinct product category alongside general reasoning models. No fresh 3.5 Flash cadence today; the news beat is the corpus-level framing hardening around the earlier ship rather than a new release.
  • 2026-08-11-AI-DigestGemini 3.5 Flash Cyber lands as Google‘s Aug leg of the three-lab cyber triopoly that crystallises with OpenAI‘s Aug 10 GPT-5.6-Cyber launch under Daybreak Red and Anthropic‘s Claude Mythos 5 under Project Glasswing — three of the four US frontier labs now shipping purpose-built cyber SKUs to gated enterprise defenders within a four-month window. Corpus framing to carry: each vendor’s Aug cyber SKU has different names and positioning (Claude Mythos 5 is Anthropic’s; Gemini 3.5 Flash Cyber is Google’s gov-and-trusted-partner variant under the AI Threat Defense umbrella); the “OpenAI joins the club” framing flattens meaningful design differences. No fresh 3.5 Flash Cyber cadence today; the news beat is the triopoly-framing anchor around the earlier ship — Meta remains the outlier and whether it ships a cyber-tuned Llama variant is the four-lab-vs-three-lab question for the next quarter.
  1. Third Leg of the Three-Lab Cyber Triopoly Alongside GPT-5.6-Cyber and Claude Mythos 5 Inside a Four-Month Window (August 11, 2026): The Jul 21 Gemini 3.5 Flash Cyber ship now reads as the earliest of the three US frontier-lab cyber SKUs that crystallise into a real triopoly with GPT-5.6-Cyber under Daybreak Red on Aug 10 — Claude Mythos 5 under Project Glasswing is the third leg. Load-bearing corpus framing: the “OpenAI joins Anthropic” framing is one lab behind — the accurate frame is three labs, three purpose-built cyber SKUs, three different names + gating shapes (Google under AI Threat Defense for gov-and-trusted-partners, Anthropic under Project Glasswing, OpenAI under Daybreak Blue/Red). Meta the outlier; four-lab-vs-three-lab is the next-quarter question.

Key Developments

  1. Flash-Tier Reframed for Agentic Workloads: Google’s explicit positioning shifts the Flash tier from “cost-optimized chat” to “long-horizon agent runs.” If the Terminal-Bench delta over 3.1 Pro survives independent evaluation, this is the first major frontier-lab Flash-tier model marketed primarily on coding-agent throughput rather than per-token cost.

  2. Treat the Benchmarks as Google-Controlled: Every datapoint at launch is on Google’s own evaluation harness with Google’s own selection of comparators (no Pro 3.5 number, no third-party replication yet). The honest read holds the numbers as directional until independent benchmarks land.

  3. Pricing Anchors but Does Not Reset the Floor: Cursor Composer 2.5 at $0.50 / $2.50 remains the cheaper agentic option on input pricing. The frontier-lab Flash tier is now in the same order of magnitude as IDE-vendor in-house models — a real convergence, but not a price floor break.