Google published version 3.1 of its Frontier Safety Framework on April 17, 2026, a 20-page document setting out how the company detects, secures and decides whether to release AI models whose abilities could contribute to severe harm. The revision adds a second, lower class of warning thresholds, merges misalignment into its machine learning research domain and hardens the protection recommended for models that cross its weapons, cyber and manipulation thresholds. More than five months later it remains the newest edition listed by Google DeepMind, and it governs the same Gemini family that now generates ad copy inside Google's advertising products.
In Short
Google keeps a rulebook for its most powerful AI models that says how it checks them for serious dangers, such as helping someone build a weapon, run a cyber attack or quietly work around human control, and what locks it puts on them. The April 2026 edition adds earlier warning lines and stricter protection for the model files, which matters to advertisers and publishers because the same Gemini family writes ad copy and answers questions in Google Search. Nothing changes inside an ad account, but the rulebook decides which model versions are released, to whom, and what governments may be told.
Six changes, four editions
According to the document's change log, version 3.1 makes six changes. It creates tracked capability levels for chemical, biological, radiological and nuclear (CBRN) risk, with a mitigation and risk acceptance process attached. The change log uses the plural, although the body of the document defines a single CBRN tracked level. It absorbs the previously exploratory misalignment domain into a combined domain covering machine learning research and development (ML R&D) and misalignment. And it lifts the recommended security for the CBRN, cyber and harmful manipulation critical capability levels to what Google calls Security Level 2+, a step the change log ties to protection "against non-state actors and insider threats." The remaining three are editorial: more detail on the risk management process, a description of internal governance and a glossary.
Version 1.0 appeared on May 17, 2024, version 2.0 on February 4, 2025 and version 3.0 on September 22, 2025, so version 3.1 makes four editions in 23 months.
Scope is narrow by design. According to Google, the framework "focuses on possible severe risks stemming from high-impact capabilities of frontier AI models," leaving other risks to the company's AI Principles. The glossary defines those models as ones whose agentic and reasoning-based general capabilities "near or exceed those of other Google models." Checkpoints and versions count as the same model when their base capabilities stem from the same foundational training run. A footnote lists the framework's peers, among them OpenAI's Preparedness Framework, Anthropic's responsible scaling policy updates and Anthropic's compliance framework for California's SB 53.
Red lines, and a new tier beneath them
The framework is built mainly around Critical Capability Levels, or CCLs. According to Google, these are "capability levels at which, absent mitigation measures, frontier AI models or systems may pose heightened risk of severe harm," each defined as the minimal set of abilities needed to cause that harm along a foreseeable path.
Version 3.1 adds Tracked Capability Levels, or TCLs, set lower to capture risk of "significant but not severe levels of harm," with mitigation applied proportionately. A third marker, the alert threshold, sits marginally below each CCL and signals, according to the glossary, that a CCL "may be reached in the foreseeable future," possibly before the next scheduled assessment.
Five CCLs and two TCLs are defined. CBRN uplift level 1 covers a model that gives low to medium resourced actors uplift producing additional expected harm at severe scale; its tracked counterpart covers the same actors where attacks could cause significant but not severe harm. Cyber uplift level 1 covers uplift with high impact cyber attacks at severe scale. Harmful manipulation level 1, labelled exploratory, covers manipulative capabilities sufficient to "systematically and substantially change beliefs and behavior in identified high stakes contexts." A footnote clarifies that uplift is measured "relative to a baseline without generative AI" at all.
The ML R&D thresholds describe a different kind of risk. Acceleration level 1 applies to a model that "has been used to accelerate AI development, resulting in AI progress substantially accelerating from historical rates." Automation level 1 goes further, to a model that "can fully automate the work of any team of researchers at Google focused on improving AI capabilities, with approximately comparable all-inclusive costs." The second tracked level, Stealth and Situational Awareness, belongs to this domain and is covered below.
From checkpoint to release
The risk management process has five stages: risk identification, inherent risk assessment, risk mitigation, residual risk assessment and risk acceptance determination. It runs, according to Google, "on various checkpoints or versions of a model, both before and after deployment."
A full critical capability assessment is required before the first external deployment of a new frontier model. Later versions trigger another only if two conditions hold: the model has meaningful new capabilities or material performance gains, and Google believes that increase "could materially undermine the justification" for accepting its risks. To decide, Google runs material capability change assessments on checkpoints after each post-training run, drawing on provenance, size, modality and general benchmark results. The early warning evaluations at the core of the process are run with "appropriate scaffolding, inference compute, and other augmentations" to approximate systems likely to be built on the model, and may be run more often if progress erodes the safety buffer.
There is an exemption. Footnote 4 states that a critical capability assessment "may not be conducted for low-risk external deployments (e.g. to a small number of trusted testers)" when the relevant governance function judges the residual risk acceptable, even for a model that has reached a threshold. For ML R&D, Google may lean on its own internal progress data, reasoning that "as a frontier AI company, we do not expect other groups to put significantly more effort into ML R&D than we do ourselves."
The versioning rules matter because releases now arrive quickly. AI Mode moved to Gemini 3.5 Flash at Google I/O on May 19, 2026. Gemini 3.7 Flash reached AI Mode for AI Pro and Ultra subscribers in English in August, and AI Mode moved to Gemini 3.8 Flash on September 2. Google DeepMind's frontier safety page listed Frontier Safety Framework reports for two models, Gemini 3 Pro and Gemini 3.7 Flash, as of September 23. The framework contains no commitment to publish such reports; its disclosure section addresses government authorities.
Security, indexed to RAND
Security mitigations are built on Google's Secure AI Framework, known as SAIF, which rests on six elements according to the document: strong security foundations, detection and response, automated defences, harmonised platform-level controls, adaptable controls for emerging threats and end-to-end risk assessments. Levels are pegged to the RAND model weight security framework, so "security level N" denotes controls generally aligned with RAND's SL N. Google says it adopts RAND's goals and principles rather than the report's concrete benchmarks.
Security Level 2+ takes RAND's SL2 as its baseline and adds measures aimed at insiders and well-resourced non-state attackers. According to the framework these may include dedicated insider risk teams, background checks and ID verification for staff with sensitive access, review of training data for signs of tampering, processing of untrusted inputs inside sandboxed environments, red-teaming that simulates well-resourced adversaries including Advanced Persistent Threat groups, and proactive threat hunting with 24/7 incident response. Untrusted input is the channel through which prompt injection reaches models that read documents and web pages on a user's behalf.
The rationale differs by domain. For CBRN, leaked weights "could be highly damaging," but the actors most likely to benefit are, according to Google, unlikely to pose an exfiltration threat at the level of RAND's OC3 groups. For cyber, such models "might be of interest to well-resourced state actors," yet automated cyber-defence and social adaptation mean "higher levels of security, and the resulting costs to innovation, are likely not warranted." For harmful manipulation, the "lower velocity of harm scenarios" counts against expensive protection.
The ML R&D levels sit higher. Acceleration level 1 carries Security Level 3, which may add measures to prevent unilateral access, harden infrastructure and prevent data exfiltration. Automation level 1 carries Security Level 4, which may add isolation of model weights, enhanced data centre security and a reduced attack surface. That recommendation arrives with a condition printed in bold: it "must be taken on by the frontier AI field as a whole." Unrestricted access at that level, the document says, "could be catastrophic if there is no effective way of defending against rapidly improving and potentially superhuman AI systems wielded by threat actors."
Insider misuse is treated as already covered: according to the framework, "baseline controls (such as background checks and access restrictions)" address it, and additional mitigations "are not required" at this stage.
The tracked levels carry no recommended security tier of their own; the document's tables cover CCLs only. The recommended levels are also adjustable. Google says they may change if a model's capabilities are not meaningfully different from other publicly available models with weaker security, or if it judges that "the benefits of the open release of model weights outweigh the risks." That clause sits alongside a public position: Google was among the 77 signatories of a July 24, 2026 letter defending open-weights models, and Google DeepMind publishes an open model family, Gemma.
The marginal-risk test
For models that reach a misuse threshold, external deployment follows three steps: developing and assessing safeguards such as safety post-training, account moderation, jailbreak detection and user verification; a pre-deployment review of residual risk by "the appropriate governance function"; and post-deployment updates driven by monitoring. Where a CCL is reached, the residual risk assessment is supplemented with a safety case, which the glossary defines as "an assessable argument showing how severe risks associated with a model's CCLs have been reduced to an acceptable level."
Five factors feed that judgement: the effectiveness of mitigations, the likelihood and consequences of misuse or circumvention, the scope of the deployment, the capabilities and mitigations of other publicly available models, and the historical incidence of related events. The fourth matters most for competition. The risk acceptance section states the principle in parentheses: "if other models are similarly capable and have few mitigations, then the marginal risk added by our external deployment is likely low."
Security adequacy follows comparable logic. For misuse and ML R&D thresholds alike, Google may treat protection as sufficient where existing mitigations "match or exceed the level of security applied to other models with similar capabilities or risk profiles." The comparison with rival products is absent, however, from the deployment factors for ML R&D and misalignment, which rely instead on model propensity and the severity of related events.
The overview frames the dependency openly. Certain mitigations are "most effective when adopted by industry as a whole," and Google's adoption of them "would result in effective risk mitigation for society only if all relevant organizations provide similar levels of protection." What happens if the rest of the field does not follow? The document does not describe a response.
Misalignment joins ML R&D
Version 3.1 treats misalignment as a tiered problem inside the ML R&D domain. The Stealth and Situational Awareness TCL marks the baseline above which misalignment risks "may begin to arise." According to the framework, it is reached when a model's instrumental reasoning gives it enough situational awareness, the "ability to discover and use relevant details of its deployment setting," and enough stealth, the "ability to circumvent basic oversight mechanisms," that "absent additional mitigations, we cannot rule out the model significantly undermining human control."
A model at that level triggers periodic residual risk assessments. Because misalignment can surface through internal use, these extend to high-risk internal deployments, which the glossary describes as releases to Google employees with potential to enable severe threat scenarios, such as building internal security infrastructure or automating ML R&D. If that internal risk is judged unacceptable without further safeguards, measures such as chain-of-thought monitoring on high-risk deployments will be applied. Other safeguards may include limiting affordances, auditing and alignment training. Risk acceptance here covers external and high-risk internal deployment, "not further development."
Governance in three sentences
Section 4, which the change log presents as new, runs to three sentences. It describes "a well-established and comprehensive internal governance structure," responsibilities "clearly defined and allocated across all levels of the organization," and "legal, compliance, and safety reviews with escalation procedures to ensure appropriate oversight." No committee, officer or reporting line is named; elsewhere the text refers only to "the appropriate governance function."
Review happens at least once a year, or sooner if Google has "reasonable grounds" to believe the framework's adequacy, or its adherence to it, has been materially undermined. Disclosure is conditional. If a model reaches a CCL that poses "an unmitigated and material risk to overall public safety," Google says it aims to share model information, evaluation results and mitigation plans with appropriate government authorities, subject to confidentiality considerations.
The language tracks that discretion. Across roughly 8,000 words, the phrase "we may" appears 10 times and "we will" eight times. The word "must" appears three times: once in a definition, once in a methodological note and once in the call on the whole frontier AI field. The closing note on risk acceptance concedes: "Because the science of AI risk assessment is still developing, our assessments will often involve some level of subjective analysis."
Why the framework reaches advertising
Why would a document about bioweapons and self-improving AI concern people who buy search ads? Because it governs models advertisers already use. Alphabet told investors in February 2026 that Gemini generates millions of creative assets through text customisation in AI Max and Performance Max. At Google Marketing Live in May, product leaders said new AI Mode ad formats write copy at query time and require AI Max or Performance Max with text customisation.
The framework already shapes decisions visible to the market. When Google released Gemini 3.6 Flash in July, it said the model carried enhanced Frontier Safety safeguards for CBRN misuse and cyber offence, while a separate Flash Cyber model was restricted to governments and trusted partners. Both safety claims were asserted with reference to a model card rather than published evaluation data. The restricted release is consistent with the framework's deployment-scope factor, under which "small scale and private deployments may pose substantially less risk than large scale or public deployments."
The harmful manipulation threshold sits closest to the industry's own trade. Persuasion is what advertising sells. The CCL is not aimed at it: it concerns misuse of models whose manipulative capabilities could reasonably produce harm "at severe scale" in "identified high stakes contexts," and the document does not say which contexts have been identified. It also marks the limits of its own knowledge, stating that "the research into harmful manipulation from a severe risk perspective is nascent." European law draws a related line. The Commission's February 2025 guidance on prohibited practices separated permissible persuasion from manipulation under Article 5 of the AI Act, a boundary adjacent to existing rules on dark patterns in interface design.
Regulators want to read these documents
Frameworks like this one are moving from voluntary publications into regulatory files. The EU's General-Purpose AI Code of Practice contains a Safety and Security chapter applying to models trained with more than 10^25 floating-point operations, and Google committed to sign the code on July 30, 2025. In California, SB 53, signed on September 29, 2025, took effect on January 1, 2026; it covers foundation models trained with more than 10^26 operations and carries penalties of up to $1 million per violation. On September 18, 2026, Governor Gavin Newsom signed Executive Order N-9-26, giving a state agency until November 16 to recommend on independent verification of frontier developers' safety frameworks, onsite verifiers inside the labs and continuing checks on a shutdown mechanism. The Stealth and Situational Awareness TCL, with its reference to models undermining human control, deals with a related scenario from inside a company.
Washington has moved on a separate track. In May 2026 the administration agreed pre-release assessment arrangements with Google DeepMind, Microsoft and xAI, and on June 2 President Donald Trump signed an order creating a voluntary review of covered frontier models up to 30 days before release. Anthropic had argued in July 2025 for requiring the largest developers to publish such frameworks, above thresholds of about $100 million in annual revenue or $1 billion in yearly R&D or capital spending. Google's own document commits only to aim to share information with governments once a CCL presents an unmitigated risk to public safety.
A document to compare, not copy
The framework has also found an audience among governance practitioners. Ksenia Laputko, Chief AI Officer and Global Head of Data Protection at Joblio, added it to a LinkedIn series that had already covered the meaning of frontier AI, the proposed US federal frontier AI governance framework and OpenAI's approach. She presents comparison as a way to see how organisations define thresholds, approach evaluations and connect technical risk management with decision-making, and she is sceptical of shortcuts: "I am not a big fan of starting from a blank page and asking an AI chatbot to simply 'write me an AI governance framework.'" According to Laputko, "good governance documents should come from understanding the field, comparing existing approaches, reading what experienced teams are already doing, and developing your own professional judgment."
Read that way, version 3.1 is a set of choices rather than a standard: where to draw thresholds, how much security to buy, and how far to benchmark against rivals. The framework commits to a review at least once a year, and California's recommendations are due on November 16.
Timeline
- May 17, 2024 - Google publishes version 1.0 of the Frontier Safety Framework
- February 4, 2025 - Version 2.0 of the framework is published
- February 2025 - European Commission guidance separates permissible persuasion from manipulation under Article 5 of the AI Act
- July 7, 2025 - Anthropic proposes requiring the largest AI developers to publish safety frameworks
- July 10, 2025 - European Commission receives the final General-Purpose AI Code of Practice, with a Safety and Security chapter for models above 10^25 FLOP
- July 30, 2025 - Google commits to sign the EU General-Purpose AI Code of Practice
- September 22, 2025 - Version 3.0 of the framework is published
- September 29, 2025 - Governor Gavin Newsom signs SB 53, the Transparency in Frontier Artificial Intelligence Act
- January 1, 2026 - SB 53 takes effect in California
- February 2026 - Alphabet says Gemini generates millions of creative assets in AI Max and Performance Max
- April 17, 2026 - Google publishes version 3.1, adding tracked capability levels, merging misalignment into ML R&D and setting Security Level 2+ for CBRN, cyber and harmful manipulation CCLs
- May 2026 - The Trump administration agrees pre-release assessment arrangements with Google DeepMind, Microsoft and xAI
- May 2026 - Google says new AI Mode ad formats write copy at query time and require AI Max or Performance Max
- June 2, 2026 - President Trump signs an order creating a voluntary pre-release review of covered frontier models
- July 2026 - Gemini 3.6 Flash ships with enhanced Frontier Safety safeguards as Flash Cyber stays restricted to governments and trusted partners
- July 24, 2026 - Google joins 77 signatories of a letter defending open-weights models
- August 2026 - Gemini 3.7 Flash reaches AI Mode for AI Pro and Ultra subscribers in English
- September 2, 2026 - AI Mode moves to Gemini 3.8 Flash
- September 18, 2026 - Newsom signs Executive Order N-9-26 on verification of frontier AI safety frameworks
- September 2026 - Ksenia Laputko adds Google's framework to her LinkedIn series comparing frontier AI governance approaches
- November 16, 2026 - Deadline for California's Government Operations Agency recommendations
Related PPC Land coverage
- Newsom sets November 16 deadline to study frontier AI kill switch - Details Executive Order N-9-26 and the California push to verify the safety frameworks frontier developers file.
- Trump signs AI order reviving the safety review he abolished 17 months ago - Covers the June 2026 federal order and the pre-release assessment arrangements that preceded it.
- Anthropic proposes new transparency framework for frontier AI models - Sets out the revenue and spending thresholds Anthropic proposed for mandatory publication of safety frameworks.
- Anthropic faces open-weights ban accusations as 77 firms sign letter - Reports the July 2026 open-weights letter that Google signed and the testing regime Anthropic proposed instead.
- Google cuts Gemini Flash prices as 3.6 uses 17% fewer output tokens - Documents the Frontier Safety safeguards claimed for Gemini 3.6 Flash and the restricted Flash Cyber release.
- EU publishes final General-Purpose AI Code of Practice - Explains the code's three chapters, including the safety obligations for systemic-risk models.
- Google commits to EU AI code amid compliance concerns - Records Google's July 2025 decision to sign the general-purpose AI code.
- EU clarifies boundary between influence and manipulation under AI Act - Examines where the AI Act draws the line between persuasion and prohibited manipulation.
- Google GML backstage: product leaders on Gemini, AI Max, and measurement - Describes how Gemini generates ad copy at query time in the new AI Mode formats.
- Alphabet advertising revenues climb 14% as Gemini App reaches 750 million users - Includes Alphabet's statement that Gemini produces millions of creative assets in AI Max and Performance Max.
- Google pushes budget-capped campaigns back up to target CPA - Notes the August 2026 arrival of Gemini 3.7 Flash in AI Mode.
Summary
Who: Google, publisher of the Frontier Safety Framework, with Google DeepMind hosting the document and its model reports. Advertisers and publishers are affected indirectly through Gemini models used in AI Max, Performance Max and AI Mode. Regulators in the EU, California and Washington, and governance practitioners such as Ksenia Laputko of Joblio, are reading and comparing the document.
What: Version 3.1 of the framework, a 20-page protocol defining five Critical Capability Levels and two Tracked Capability Levels across CBRN, cyber, harmful manipulation and ML R&D and misalignment risk. It sets Security Level 2+ for the CBRN, cyber and harmful manipulation CCLs, Security Level 3 and 4 for the two ML R&D CCLs, a five-stage risk management process, a residual risk test that weighs rival models' capabilities and mitigations, and conditional disclosure to governments.
When: Published on April 17, 2026, following versions dated May 17, 2024, February 4, 2025 and September 22, 2025. It remained the latest version listed by Google DeepMind on September 23, 2026, with annual review promised and California recommendations on frontier safety frameworks due on November 16, 2026.
Where: The framework applies to Google's frontier models across internal and external deployments worldwide, including Gemini models serving Google Search and Google Ads. The regulatory context spans the EU's General-Purpose AI Code of Practice, California's SB 53 and Executive Order N-9-26, and the US federal pre-release review.
Why: The framework decides which model versions Google releases, to whom and with what protection, and those models now write ad copy and search answers at scale. Its reliance on company discretion, its comparison with rivals' safeguards and its admission of subjective judgement are the kind of self-assessment California is now studying whether to verify independently.
Discussion