Google released three Gemini models on July 21, 2026, cutting the price of its highest-volume Flash tier while restricting a specialised cybersecurity variant to a closed pilot programme. The company also disclosed that pre-training has begun on Gemini 4.
The announcement, published on Google's Keyword blog by Tulsee Doshi, Senior Director of Product Management, on behalf of the Gemini team, covers Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. Two of the three shipped immediately. The third did not.
For marketers, the numbers that matter are not the benchmark scores. They are the token counts and the per-million prices, because those determine what an automated workflow costs to run at volume.
Token efficiency becomes the pricing story
According to Google, Gemini 3.6 Flash consumes 17% fewer output tokens than its predecessor, 3.5 Flash, as measured on the Artificial Analysis Index. On DeepSWE by Datacurve, a software engineering benchmark, the company observes reductions reaching 65%.
That gap between 17% and 65% is wide, and Google's phrasing is careful. The 17% figure is described as an index-wide measurement. The 65% figure applies to specific benchmarks under specific conditions. Anyone budgeting on the higher number is budgeting on a best case.
Price moved alongside efficiency. Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens, which Google states is lower than 3.5 Flash. The company frames the combined effect as a reduction in cost per agentic task rather than a simple rate cut, since fewer tokens at a lower rate compounds.
Google also states the model takes fewer reasoning steps and tool calls to complete multi-step workflows. For agent operators, step count is a cost driver independent of token price, since each tool call carries its own latency and, frequently, its own third-party charge.
Benchmark movement
Google published paired comparisons against 3.5 Flash across several evaluations:
- DeepSWE: 49% against 37%
- MLE Bench, covering machine learning research tasks: 63.9% against 49.7%
- OSWorld-Verified, covering computer use: 83.0% against 78.4%
- GDPval-AA v2, covering knowledge work: 1421 against 1349
The company attributes the DeepSWE improvement to higher precision, with fewer unwanted code edits and reduced execution loops. Execution loops are a recognised failure mode in autonomous agents, where a model repeats an action without converging on a result, burning tokens with no output.
Computer use is now a built-in client-side tool in the Gemini API and Gemini Enterprise, according to Google. That capability was previously introduced in Gemini 3.5 Flash as a distinct feature.
Google named Hebbia and Harvey as customers finding 3.6 Flash capable at multimodal tasks including document parsing, chart and data analysis, and report drafting. No specific performance figures accompanied those customer references.
Flash-Lite targets throughput
Gemini 3.5 Flash-Lite occupies a different position. According to Google, it runs at 350 output tokens per second as measured by Artificial Analysis, making it the fastest model in the 3.5 series. Pricing sits at $0.30 per million input tokens and $2.50 per million output tokens.
Those figures put Flash-Lite output at one third the price of 3.6 Flash output. For high-volume, low-complexity work - the kind of repetitive classification and extraction that sits underneath most production advertising automation - that difference determines whether a workflow is economically viable.
Google positions the model for agentic search, document processing, and other tasks where throughput rather than depth is the constraint. The comparisons published are against 3.1 Flash-Lite rather than 3.5 Flash:
- Terminal-Bench 2.1: 54% against 31%
- GDM-MRCR v2, covering long context: 72.2% against 60.1%
- GDPval-AA v2: 1140 against 642
The GDPval-AA v2 jump from 642 to 1140 is the largest proportional gain Google disclosed for any model in the release. It still leaves Flash-Lite well below the 1421 recorded by 3.6 Flash on the same benchmark, which is the trade-off the tiering is designed to express.
More notable for anyone running existing deployments: Google states that Flash-Lite outperforms Gemini 3 Flash on several agentic and coding evaluations, citing SWE-Bench Pro at 54.2% against 49.6% and OSWorld-Verified at 74.0% against 65.1%. If accurate, that means workloads currently running on 2.5 Flash or 3 Flash face a cheaper and faster option in a model Google classifies as a lightweight tier.
Developers can configure thinking levels, according to Google, selecting minimal or low settings to prioritise latency and cost on high-volume tasks, or engaging higher levels for multi-step subagent workloads. Computer use is a built-in tool here as well.
Cyber model stays behind a gate
The third model, Gemini 3.5 Flash Cyber, did not ship publicly and is the most consequential item in the release for reasons unrelated to advertising.
Google's framing is direct. AI models have become capable of finding security vulnerabilities faster than current systems can fix them, according to the company. Flash Cyber is built on 3.5 Flash and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models.
It operates inside CodeMender, Google's code security agent, which the company states runs multiple Flash Cyber agents in parallel to produce a single combined report. Google says the combination reaches competitive performance at the frontier on CyberGym, a vulnerability benchmark.
Distribution is where the announcement diverges from the rest. Citing the dual-use nature of the technology, Google stated the model will be exclusively available to governments and trusted partners via CodeMender as part of a limited-access pilot programme, with the stated aim of giving defenders a head start while mitigating broader misuse. No date was attached to that availability beyond the word soon.
The same tension appears in the safety disclosure for 3.6 Flash. Google states the model ships with enhanced Frontier Safety safeguards covering chemical, biological, radiological, and nuclear misuse alongside cyber offence, describing it as substantially more resistant to jailbreaks while trained to minimise refusals for beneficial uses. Both claims are asserted; the release points to a model card rather than published evaluation data.
Availability and what is missing
Gemini 3.6 Flash and 3.5 Flash-Lite are available immediately across several surfaces, according to Google:
- The Gemini API via Google AI Studio and Android Studio, with 3.6 Flash also in Google Antigravity
- Gemini Enterprise Agent Platform, with 3.6 Flash also in the Gemini Enterprise app
- The Gemini app for all users, with 3.5 Flash-Lite also rolling out in Google Search
That last line carries the most weight for search marketers and publishers. Flash-Lite is entering Google Search. The company did not specify which Search surfaces, which markets, what proportion of queries, or over what period.
Gemini 3.5 Pro remains in partner testing, with Google stating it plans broad availability as soon as it is ready. And the company disclosed that its most ambitious pre-training run yet has started, for Gemini 4.
Why the numbers matter for marketing operations
The rollout of Flash-Lite into Search continues a pattern PPC Land has tracked through 2026. At I/O 2026, Google upgraded AI Mode to Gemini 3.5 Flash as its default model globally, replacing the prior model across every country where AI Mode operates. That followed the January 27, 2026 change making Gemini 3 the default for AI Overviews worldwide.
Each swap has moved in the same direction: toward models optimised for speed and unit cost rather than maximum capability. The reason is arithmetic. AI Mode surpassed one billion monthly active users, and its queries run roughly three times as long as conventional search queries, according to figures Google published alongside the I/O announcements. Serving that volume at frontier-model prices is not sustainable, which is why Alphabet disclosed that hardware and engineering work had reduced the cost of core AI responses by more than 30% since Gemini 3 launched, a figure it presented while raising approximately $85 billion for AI infrastructure and revising 2026 capital expenditure guidance to between $180 billion and $190 billion.
Cheaper inference expands the set of queries that can economically receive an AI response. That expansion is the mechanism through which AI surfaces have displaced conventional result formats, and the collapse in position-one click-through rates from 27% to 11% documented by SISTRIX in March 2026 is the visible consequence for publishers.
The agent economics carry a separate implication. Google's own advertising infrastructure now depends on Gemini across fraud detection, campaign optimisation, image generation, and search products, and Alphabet reported in February 2026 that Gemini processes over 10 billion tokens per minute through direct API use, with the Gemini app reaching 750 million monthly active users. A 17% reduction in output tokens against that baseline is a material change in serving cost, whichever party is paying.
For agencies and in-house teams building automation on the Gemini API, the tier restructuring changes the calculation on which model handles which task. Google has now published a lightweight model that beats a previous-generation mid-tier model on coding and agentic benchmarks while costing less per token. Whether the published gains hold outside Google's chosen benchmarks is a question the release does not answer, and the customer references in the announcement carry no measurements.
There is also a governance dimension that has surfaced repeatedly in Gemini API coverage. Google added project-level spend caps to the Gemini API in March 2026 following billing incidents that charged developers for output tokens they had not generated, and separate research disclosed in February 2026 found more than 2,800 live Google Cloud API keys that had silently gained the ability to authenticate against Gemini endpoints. Faster and cheaper models increase the volume of automated traffic flowing through that same credential architecture.
The Flash Cyber restriction, meanwhile, is a rare instance of a major platform withholding a shipped capability on stated misuse grounds while simultaneously publishing its benchmark performance. The model exists, performs at a level Google describes as competitive at the frontier, and is not being sold.
Timeline
- November 2025 - Google launches Gemini 3 with generative UI capabilities for dynamic search experiences
- January 27, 2026 - Gemini 3 becomes the default model for AI Overviews globally, announced by Robby Stein, VP of Product for Google Search
- February 2026 - Alphabet reports Gemini processing over 10 billion tokens per minute via direct API use, with the Gemini app at 750 million monthly active users
- March 2026 - SISTRIX data shows position-one click-through rates falling from 27% to 11%
- March 2026 - Google adds Project Spend Caps to the Gemini API after developer billing incidents
- May 19, 2026 - Google I/O 2026 introduces the Gemini 3.5 series and Antigravity 2.0; AI Mode is upgraded to Gemini 3.5 Flash globally and surpasses one billion monthly active users
- June 3, 2026 - Alphabet discloses an approximately $85 billion equity raise and revises 2026 capital expenditure guidance to between $180 billion and $190 billion
- June 8, 2026 - NotebookLM upgrades to Gemini 3.5 and adds a secure cloud computer with 11 downloadable output formats
- July 21, 2026 - Google publishes Gemini 3.6 Flash and 3.5 Flash-Lite for immediate availability, introduces 3.5 Flash Cyber inside CodeMender as a limited-access pilot, and confirms pre-training has begun on Gemini 4
Related PPC Land coverage
- Google Search gets agents, a new search box, and Gemini 3.5 at I/O 2026 - Documents the May 19, 2026 replacement of the AI Mode default model with Gemini 3.5 Flash across nearly 200 countries.
- Gemini 3.5 and Antigravity 2.0 headline Google I/O 2026 reveal - Covers the developer keynote where the Gemini 3.5 series and the WebMCP proposal were introduced.
- Google's AI Overviews upgrade: Gemini 3 powers smoother chat handoffs - Reports the January 27, 2026 decision to make Gemini 3 the global default for AI Overviews.
- Alphabet raises $85 billion to bet everything on AI infrastructure - Details the capital raise and the disclosed 30% reduction in core AI response costs since Gemini 3 launched.
- Google finally adds Gemini API spend caps - after billing chaos hit devs - Traces the billing failures that preceded project-level spending limits in the Gemini API.
- Google API keys weren't secrets - until Gemini changed the rules - Examines the credential exposure affecting more than 2,800 live keys and the token volumes flowing through the Gemini API.
- Google I/O 2026: the search changes SEOs did not expect - Sets out the click-through rate data and search interface changes that followed the model swap.
- Inside Google I/O 2026: the agentic AI shift no one saw coming - Records the technical rationale Google leaders gave for the Flash architecture and agentic direction.
- Google lets Gemini agents run background tasks without breaking connections - Covers credential handling inside long-running managed agent sessions on the Gemini API.
- NotebookLM gets Gemini 3.5, a cloud computer, and 11 output formats - Documents how Gemini 3.5 was deployed into a standalone research product ahead of the 3.6 release.
Summary
Who: Google, through Tulsee Doshi, Senior Director of Product Management, writing on behalf of the Gemini team. The release affects developers building on the Gemini API, enterprises using the Gemini Enterprise Agent Platform, Gemini app users, and, through the Flash-Lite rollout into Search, marketers and publishers dependent on Google Search traffic. Governments and unnamed trusted partners are the only parties with access to the cyber model.
What: Three models. Gemini 3.6 Flash, priced at $1.50 per million input tokens and $7.50 per million output tokens, consuming 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index and up to 65% fewer on DeepSWE. Gemini 3.5 Flash-Lite, priced at $0.30 and $2.50 per million tokens and running at 350 output tokens per second, which Google states outperforms Gemini 3 Flash on SWE-Bench Pro and OSWorld-Verified. And Gemini 3.5 Flash Cyber, a vulnerability-focused model deployed inside the CodeMender agent and withheld from general release. Google also confirmed that pre-training has started on Gemini 4 and that Gemini 3.5 Pro remains in partner testing.
When: July 21, 2026. Gemini 3.6 Flash and 3.5 Flash-Lite became available the same day. Flash Cyber has no announced availability date beyond a limited-access pilot described as arriving soon.
Where: The Gemini API via Google AI Studio and Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform and Gemini Enterprise app, the Gemini app, and Google Search, where Flash-Lite is rolling out without disclosed scope or market detail.
Why: Google states that developers and customers building production AI agents require higher token efficiency, lower latency, and more reliable performance, and that the Flash series is built to meet that combination at scale. The commercial logic is unit economics: serving AI responses across a Search surface exceeding one billion monthly active users depends on driving inference cost down, and each successive Flash generation has been positioned as cheaper and faster rather than more capable in absolute terms.
Discussion