On July 21, 2026, Google released three Gemini models at once. Two of them, Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, went out the way models normally do: available to anyone with an API key. The third did not. Gemini 3.5 Flash Cyber, a version of Flash fine-tuned specifically to find and patch software vulnerabilities, is "exclusively available to governments and trusted partners" through Google's CodeMender program, according to Google DeepMind's own announcement. There is no general-availability date, no public pricing, and no path for an ordinary developer to use it.
That single product decision is not an outlier. It is the clearest recent instance of a pattern that now spans four major AI companies: the most capable, most security-relevant version of a model is no longer released to everyone. It is released to a vetted list, sometimes with a government sitting between the lab and its customers. This piece looks at what that pattern actually consists of, why cyber capability specifically is what triggers it, how it compares to older export-control regimes, and what it means for the much larger group of users who will never touch a gated model at all.
AnthropicWhat Google shipped, and who is allowed to use it
Gemini 3.5 Flash Cyber is built on the same Flash architecture as the two models released alongside it, but fine-tuned for vulnerability discovery, validation and patching. Google says it deploys the model through CodeMender, its AI vulnerability-fixing agent first unveiled in October 2025, which calls the smaller Cyber model repeatedly and cheaply to scan far more code paths than a single expensive frontier-model pass would allow. In a benchmark against the V8 JavaScript engine, Google reports the Cyber variant found 55 confirmed unique vulnerabilities, against 47 for standard Gemini 3.5 Flash and 36 for Anthropic's Claude Opus 4.6.
Why Google says the gated variant finds more bugs
Confirmed unique vulnerabilities found in the V8 JavaScript engine benchmark. A single vendor-run test, not an independent audit.
Google's stated rationale for gating the model: the same fine-tuning that finds more real vulnerabilities faster is the capability it says needs a vetted-access pilot rather than a public launch.
Those are Google's own numbers from a single vendor-run benchmark, not an independently audited result, and should be read with that caveat. But the gap is exactly the property Google cites as its reason for restricting access in the first place: a model that is measurably better at finding real vulnerabilities faster is also, by the same logic, a model that is better at finding them for an attacker. Google's own framing states the model will "give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse," which is a direct statement of the dual-use tradeoff driving the access decision.
Google is the newest entry, not the first
The Gemini 3.5 Flash Cyber launch is notable mainly because it makes the pattern legible in a single week. Three other companies had already built essentially the same structure.
Anthropic went first, and hit the hardest version of the problem. In April 2026, it announced Mythos Preview, a model it judged too capable in offensive cyber terms for a public release, alongside Project Glasswing, a defensive-access program that shared the model with roughly 50 vetted partners, several connected to the US government. In June, Anthropic expanded Glasswing to around 150 additional organizations across more than 15 countries, spanning power, water, healthcare and communications infrastructure, and said partners had by then surfaced more than 10,000 high- or critical-severity flaws in their own codebases. Anthropic's own description of the tradeoff is direct: "because cybersecurity has both helpful and destructive uses, making safeguards that are both strong and precise enough is a major challenge."
That same month, the arrangement collided with the US government directly. On June 12, 2026, the Department of Commerce sent Anthropic an export-control directive suspending all foreign-national access to Mythos 5 and its sibling model, Fable 5, reportedly after a jailbreak vulnerability raised concern inside the administration. On June 26, the government reversed course far enough to authorize a narrower release, roughly 100 US companies and federal agencies, and by July 1 Anthropic had restored Fable 5 globally while keeping Mythos 5 limited to the newly approved US list. Whatever one makes of the specific episode, the sequence establishes something structurally new: a US federal agency acting as a distribution chokepoint for a specific frontier model's release list, not merely as a regulator of the industry in general.

OpenAI built a parallel structure under the name Daybreak, launched May 12, 2026, with an explicit three-tier design: standard GPT-5.5 for common defensive work, "GPT-5.5 with Trusted Access for Cyber" for verified defenders needing more precise safeguards, and GPT-5.5-Cyber, expanded June 22, reserved for verified defenders whose authorized work, such as red-teaming or exploit development, needs the model's most permissive and capable behavior. Microsoft is reportedly preparing a fourth version of the same idea, Project Perception, an AI security platform that itself routes across models from Microsoft, OpenAI and Anthropic, with reporting suggesting Microsoft may restrict initial access to vetted companies, government organizations and compliance-qualified customers rather than launching broadly. (Metir has covered Project Perception's multi-model routing design and Gemini 3.6 Flash's pricing changes separately; this piece is about the access-gating pattern connecting all four programs, not either product individually.)
Three tiers of access to security-tuned frontier models
Four different companies, one shared shape: the most cyber-capable variant of a model sits behind the narrowest gate.
- Gemini 3.6 Flash & 3.5 Flash-Lite (Google)
- GPT-5.5, standard tier (OpenAI)
- Claude Opus 4.8 / Claude Security (Anthropic)
- GPT-5.5 with Trusted Access for Cyber (OpenAI Daybreak)
- Claude Mythos Preview via Project Glasswing (~200 orgs, 15+ countries)
- Project Perception pilot customers (Microsoft, reported)
- Gemini 3.5 Flash Cyber, CodeMender pilot (Google)
- GPT-5.5-Cyber, verified defenders only (OpenAI Daybreak)
- Claude Mythos 5, ~100 US companies & federal agencies post export-control review (Anthropic)
Tier boundaries are each company's own characterization of its program and can shift; several (Daybreak, Glasswing) have already expanded since launch.
Why cyber capability specifically triggers this
Plenty of AI capabilities are dual-use in some loose sense. Cyber capability is different in a way that makes gating a more direct response than it would be elsewhere: finding and exploiting a software vulnerability is close to a single skill, exercised identically whether the actor is a defender patching their own system or an attacker preparing to break into someone else's. A model that gets meaningfully better at that skill does not split cleanly into a "defensive" version and an "offensive" version. It is the same capability pointed in two directions, and the direction is determined entirely by who holds the credentials, not by anything in the model itself.
That is a sharper version of dual-use than, say, a model that writes persuasive text or generates images, where misuse and legitimate use are still separable by context and downstream controls. With vulnerability discovery, the fastest way for a lab to reduce offensive uplift is to reduce who can run the query at all, because there is no way to make the underlying capability serve only defenders once it is broadly available. That is the logic every one of the four programs describes in its own words, and it is a more coherent reason for gating than most capability restrictions get to claim.
With most AI capabilities, misuse and legitimate use diverge somewhere downstream. With vulnerability discovery, they are the same query, and the only variable is who is allowed to run it.
Metir AI analysis
An export-control regime, or something new
It is tempting to call this an AI version of export controls, and the comparison holds in one important respect: in both cases, a government is deciding who outside its own borders may access a technology capable of national-security-relevant harm, and a company's commercial release plan is subordinated to that decision. The June 2026 Commerce Department directive against Anthropic's models is functionally an export-control action in exactly that sense.
But the analogy breaks in ways that matter for how durable this regime can be. Traditional export controls govern physical hardware and classified technical data: a chip, a missile component, a document that can be seized at a border or tracked through a supply chain with serial numbers. A frontier model is a set of weights that can be copied at effectively zero cost, does not degrade with duplication, and crosses no physical checkpoint. Enforcement therefore depends entirely on controlling the interface, the API, the account approval, the contractual terms, rather than the artifact itself. That is a much weaker enforcement mechanism than a border inspection, and it means the entire regime rests on the good-faith compliance of a handful of companies that could, in principle, simply choose not to comply, or could see a foreign lab release a comparable model with no gating at all. Anthropic itself has flagged this: expanding Glasswing partly reflects an expectation that competitors will ship comparable capability without equivalent safeguards within six to twelve months.
The other structural difference is who is setting the criteria. Export-control law has decades of statutory definitions, licensing schedules and judicial review behind it. The current AI gating regime is largely improvised: a voluntary executive-order framework finalized only this August, individual company vetting criteria that are not public, and approval lists whose membership and standards are disclosed, if at all, at the company's discretion. That is a meaningfully different governance foundation than the one the export-control comparison implies.
A simple way to think about which capabilities get gated
Reading across all four programs, the gating decision seems to track three factors more than any single one: how much uplift the model gives over tools an attacker already has, how costly the capability would be for a well-resourced adversary to replicate independently, and how concentrated the defensive value is (whether it mainly helps a narrow set of large, already-resourced defenders, or genuinely widens access for smaller organizations that could not otherwise afford dedicated security research). A capability that scores high on uplift and low on independent-replication cost is the one every one of these programs has chosen to gate; a capability that scores low on uplift, such as general-purpose chat or coding assistance, has not attracted the same treatment despite plenty of loose "dual use" framing applied to AI generally. That framework is a useful lens for judging the next gated release, whichever company ships it, rather than treating every restricted-access announcement as equally justified or equally alarming.
The transparency tradeoff, stated neutrally
Every one of these programs makes a real security argument, and every one of them also makes independent verification harder. A model available only to a self-selected, non-public list of organizations cannot be red-teamed by the broader research community the way a public release can, and a government-approved list is not the same thing as a democratically accountable one. Reasonable people can read the current moment two ways. One reading is that gating is simply prudent: a small number of highly capable systems are being kept away from a very large number of potential bad actors, at real cost in openness, but proportionate to a genuinely unusual risk. The other reading is that a small number of politically connected companies and a small number of government offices are quietly becoming the sole judges of who gets access to a category of tool with growing economic and security value, with no independent body checking whether the vetting criteria are consistent, fair, or even coherent across programs. Both readings are currently defensible, and the honest position is that the evidence available in July 2026 does not settle which one is closer to correct.
What this actually changes for most AI users
It is worth being precise about scale here. The organizations affected by any of these four programs number in the hundreds, out of a user base for mainstream AI models that runs into the hundreds of millions. Nearly every enterprise team, developer and individual user will only ever interact with the generally-available tier of these models: Gemini 3.6 Flash, GPT-5.5, Claude Opus 4.8, and their successors, none of which carry any gating at all. The government-only tier is a real and analytically interesting development, but it sits at the extreme edge of the model landscape, not in the middle of it.
What it does add, even for ordinary users, is one more axis of fragmentation to track: which capability lives in which tier, under which vendor's terms, changes every few months as these programs expand or contract, the same way pricing and benchmark scores shift with each release. That is a smaller version of a problem AI teams already manage by routing work across providers rather than standardizing on one model for everything. A model-agnostic workspace such as Metir AI does not touch the governed tiers this piece describes, but it reflects the same underlying idea: as access, pricing and capability keep shifting model by model and program by program, giving teams a single place to reach whichever generally-available model fits a given task is a more durable strategy than betting on any one vendor's roadmap staying still.
The open questions
Several things about this pattern remain genuinely unsettled. The voluntary federal framework behind the June 2026 executive order is not due to be finalized until August 1, 2026, so the government's role could still formalize into something closer to binding licensing or stay closer to the current ad hoc arrangement. None of the four companies discussed here has published detailed, comparable vetting criteria, so it is not possible to independently verify how consistently "trusted partner" is being defined across programs. And Anthropic's own expectation, that comparable models will appear elsewhere without equivalent safeguards within a year, has not yet been tested. Whether approval-gating becomes a stable feature of frontier AI releases or a transitional phase specific to this generation of cyber-capable models is the question the next twelve months will actually answer.
Sources:
- Introducing Gemini 3.5 Flash Cyber | Google DeepMind
- Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber | Google Blog
- Google Launches Gemini 3.5 Flash Cyber AI to Find and Fix Software Vulnerabilities | The Hacker News
- Google is taking on Anthropic's expensive Mythos model with Gemini 3.5 Flash Cyber | Chrome Unboxed
- Expanding Project Glasswing | Anthropic
- Anthropic expands Mythos to 150 additional organizations in more than 15 countries | CNBC
- Anthropic scales Claude Mythos to critical infrastructure in 15+ countries | TechCrunch
- Statement on the US government directive to suspend access to Fable 5 and Mythos 5 | Anthropic
- Trump admin allows Anthropic to release Mythos AI model to some companies, government agencies | CNBC
- Anthropic's Mythos 5 AI Model Cleared by US for Wider Use | Bloomberg
- Anthropic restores AI models Fable, Mythos after the U.S. lifts export controls | CoinDesk
- Anthropic reactivates Fable, Mythos after securing government approval | Cybersecurity Dive
- OpenAI Daybreak: Trusted Access for Cyber Overview | OpenAI Help Center
- OpenAI Expands Daybreak With GPT-5.5-Cyber to Help Defenders Patch Security Flaws | The Hacker News
- Daybreak: Tools for securing every organization in the world | OpenAI
- Microsoft's Project Perception Could Challenge Anthropic's Mythos in AI Security | TechRepublic
- Microsoft Project Perception: AI Bug Hunter Set to Rival Mythos With Wider Access and Lower Cost | Tech Times
- Trump AI executive order sets 30-day frontier model review | The Register
- President Trump Signs Executive Order Establishing AI Cybersecurity and Frontier Model Framework | Latham & Watkins
Image credits
Header image: the Google headquarters sign at 1600 Amphitheatre Parkway, Mountain View, California, photographed by Håkan Dahlström via Wikimedia Commons, licensed under CC BY 2.0. In-body photograph: the entrance of the Herbert C. Hoover Building, headquarters of the U.S. Department of Commerce in Washington, D.C., photographed by Carol M. Highsmith via the Library of Congress Carol M. Highsmith Archive on Wikimedia Commons, public domain. Neither photo depicts a specific 2026 model launch or government review; both illustrate the real organizations discussed in the piece.
