metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
Gemini 3.8 Flash
Google DeepMind
AI Models
AI Agents
Cybersecurity AI

Gemini 3.8 Flash and 3.8 Flash Cyber: The Real Numbers

Google shipped Gemini 3.8 Flash and a restricted 3.8 Flash Cyber variant on September 2, 2026. The benchmarks show a genuine split verdict, not a clean win.

Metir AI TeamSeptember 2, 20268 min read
Gemini 3.8 Flash and 3.8 Flash Cyber: The Real Numbers

On September 2, 2026, Google released two new models at once: Gemini 3.8 Flash, generally available across Google's developer and consumer surfaces, and Gemini 3.8 Flash Cyber, a security-specialist variant restricted to vetted defenders. Read the benchmark tables Google itself published, and the honest story is not "Google's new model beats the competition." It is a split verdict: 3.8 Flash clearly leads on bounded, domain-specific agent tasks, and clearly trails on the kind of open-ended computer operation that has become its own benchmark category this year.

Google logoGoogle
Gemini logoGemini
Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, released September 2, 2026.

That split is worth taking seriously rather than rounding off in either direction, because it says something about where Flash-tier models are actually improving in 2026, and where the gap to a frontier model like Anthropic's Claude Opus 5 has not moved much at all.

What actually shipped

Gemini 3.8 Flash sits in Google's mid-tier lineup, with a 1 million token context window, a 64,000 token max output, and a default thinking level of medium. It is available immediately through Google AI Studio, Android Studio, the Gemini API and Stitch for developers, through Gemini Enterprise for business customers, and through the Gemini app, AI Mode in Google Search and Google Sheets for Google AI Pro and Ultra subscribers.

Gemini 3.8 Flash Cyber is a narrower release. It is only available to "trusted defenders who require a comprehensive set of cyber capabilities" through Google's Fairwind Program, which covers governments, critical infrastructure operators and software maintainers. It is not generally available, and Google has not published a timeline for wider access.

Sep 2, 2026Release dateTwo models, one GA and one restricted
$0.75 / $3.75Intro price per 1M tokensInput / output, through Dec 31, 2026
61.4%Vals Finance Agent v2vs 58.6% for Claude Opus 5
19.1%Terminal-bench 4.0vs 51.8% for Claude Opus 5

Pricing: the same rate, a different story on tokens

Gemini 3.8 Flash carries an introductory price of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, rising to $1.50 and $7.50 on January 1, 2027. Context caching reads are $0.075 per million tokens introductory, rising to $0.15. Those are exactly the same numbers Gemini 3.7 Flash carried at its own launch, so in strict per-token terms this upgrade is free.

On bounded domain-agent tasks, 3.8 Flash leads both rivals

Google-reported pass rates at launch, September 2, 2026. Higher is better. These are finance, legal and biology agent benchmarks with a defined task and a checkable answer.

Gemini 3.8 Flash (green) tops Claude Opus 5 and GPT-5.6 Sol on all three benchmarks shown here.

But Google attached an unusual caveat to this release: 3.8 Flash "works harder" on complex tasks and can consume more tokens than 3.7 Flash did on the same job, and Google's own guidance to developers who care about cost is to use a lower thinking effort level or simply stay on 3.7 Flash. That is a real reversal of framing from just six weeks earlier. Gemini 3.6 Flash was pitched explicitly on token efficiency, using up to 17% fewer output tokens than its predecessor on the same tasks. Gemini 3.8 Flash is pitched on doing more thinking per task instead, at an identical sticker price. The result is that a flat per-token rate does not guarantee a flat, or lower, cost per completed task, and teams that track spend by the job rather than by the token should actually measure it rather than assume the price page tells the whole story.

Where 3.8 Flash wins: bounded, checkable agent tasks

On the domain-specific agent benchmarks Google published, Gemini 3.8 Flash leads both Claude Opus 5 and OpenAI's GPT-5.6 Sol by a real margin. On Vals Finance Agent v2, a benchmark of finance-specific agentic tasks, 3.8 Flash scores 61.4% against 58.6% for Opus 5 and 53.8% for GPT-5.6 Sol. On Harvey's Legal Agent benchmark, it scores 10.0% against 6.7% and 2.5% respectively, low absolute numbers that reflect how hard legal-agent tasks are across the board, but a consistent lead. On BioMysteryBench's harder "Human Difficult" tier, 3.8 Flash scores 56.5% against 49.4% for Opus 5, 44.7% for GPT-5.6 Sol, and 43.5% for its own predecessor, 3.7 Flash.

“

The model that wins the finance and legal agent benchmarks is the same model that loses badly at simply operating a computer. Both facts are true about Gemini 3.8 Flash at once.

Metir AI analysis

The pattern extends to several other benchmarks Google published without a like-for-like comparison against 3.7 Flash: CharXiv Reasoning (86.2% versus 83.7% for Opus 5), LVBench in agentic mode, a long-video reasoning test (87.8% versus 75.4%), HLE-Verified (54.9% versus 54.4%), and LABBench2, a life-sciences benchmark (86.2% versus 84.2%). Not every comparison favors 3.8 Flash: on BioMysteryBench's easier "Human Solvable" tier, Opus 5 edges ahead at 90.1% against 88.8%, and on the broad knowledge-work benchmark GDPVal-AA v2, Opus 5 scores 1824 against 1545 for 3.8 Flash, a meaningful gap.

Benchmark3.8 FlashOpus 5
Terminal-bench 2.189.4%89.1%
CharXiv Reasoning86.2%83.7%
LVBench (agentic mode)87.8%75.4%
HLE-Verified54.9%54.4%
LABBench286.2%84.2%
BioMysteryBench (Human Solvable)88.8%90.1%
GDPVal-AA v215451824

Where it loses: open-ended computer use

The clearest counterweight sits in a different family of benchmarks, the ones that test whether a model can operate a real computer or terminal across a long, unscripted task rather than answer a bounded question. Here the gap to Opus 5 is large and does not close. On Terminal-bench 4.0, Gemini 3.8 Flash scores 19.1% against 51.8% for Opus 5, a difference of more than two and a half times. On OSWorld-2.0, a desktop computer-use benchmark, it scores 59.0% against 75.4%. On DeepSWE v1.1, a software-engineering agent benchmark, it scores 71.0% against 74.0%, a smaller but still real gap.

On open-ended computer use, the gap to Opus 5 stays wide

Google-reported scores, September 2, 2026. 3.8 Flash improves over its own predecessor on both benchmarks where a 3.7 Flash score was published, but Claude Opus 5 still leads by a large margin on long-horizon terminal and desktop tasks. DeepSWE v1.1 was not published for 3.7 Flash.

Gemini 3.8 Flash (green) narrows the gap to Claude Opus 5 (gray) versus Gemini 3.7 Flash (light gray), but does not close it.

3.8 Flash does improve generationally on every one of these against 3.7 Flash, which scored 11.2% on Terminal-bench 4.0 and 50.6% on OSWorld-2.0. That is genuine progress. It just is not enough progress to close the distance to Opus 5, and the size of the remaining gap on Terminal-bench 4.0 in particular suggests Google's Flash tier is still built for shorter, more structured agent work rather than the kind of long-horizon, exploratory computer operation Opus 5 is tuned for. There is an added wrinkle worth noting: on the older Terminal-bench 2.1, 3.8 Flash actually edges ahead of Opus 5, 89.4% to 89.1%. Two benchmarks with the same name and a similar premise point in opposite directions, which is a useful reminder that a benchmark version number is doing real work, and that "computer-use benchmark" is not one settled test but a moving target that gets harder as models improve against the older versions.

Wide exterior view of Google's Gradient Canopy office building on Google's Mountain View campus
Google's Gradient Canopy building on its Mountain View, California campus. The photo shows Google's real office campus, not the Gemini 3.8 Flash launch itself.

Gemini 3.8 Flash Cyber: gated access, mixed benchmark, real deployment evidence

The Cyber variant is the more unusual release of the two, and not primarily because of its headline benchmark. On CWE-Bench, a vulnerability-detection benchmark operated by Collinear, 3.8 Flash Cyber scores 47.2% pass@1, slightly behind the 47.8% reported for the leading frontier model. Google is not claiming a benchmark win here. What it is claiming, and what its named partners appear to be validating, is deployment performance in a different sense: reliability and cost in production security workflows.

3.8 Flash Cyber reports over 70% success on Google's internal multi-language vulnerability discovery benchmark spanning 20 languages, outperforms both its own predecessor, 3.5 Flash Cyber, and larger models on CyberGym, and shows improved robustness on the Gray Swan prompt-injection benchmark. Google frames the model's priorities explicitly: it says it "invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation," a defender-first positioning that is consistent with the model being gated behind the Fairwind Program rather than sold broadly.

The deployment numbers from named partners are the more concrete evidence. Google's Chrome Security team reports the model producing 2.6 times more correct patches than commercial alternatives it tested against. Wiz reports 7.5 to 9.7% higher recall at 2.3 to 5.2 times lower cost. Google Cloud's own Vulnerability Research team found a critical vulnerability in under two hours using the model. Executives from three partner organizations, Armadin's David Slater, Palo Alto Networks' Charlie Sestito, and Snowflake's Mayank Upadhyay, are named giving testimonials in Google's announcement. A benchmark that lands just behind the frontier, paired with production evidence like that, is a genuinely different kind of pitch than a leaderboard-topping score would have been.

What this means for teams choosing a model

The practical read for teams building agents is not "switch to Gemini 3.8 Flash" or "stay on Claude." It is that the right model now depends heavily on the shape of the task. A workflow built around structured, checkable agent work, finance research, legal document review, biology or long-video analysis, has a real reason to evaluate 3.8 Flash on its own merits. A workflow built around a model driving a real desktop or terminal through a long, unscripted task still has a real reason to reach for Opus 5 instead, at least until the next Flash-tier release narrows that specific gap further. Locking a product into one vendor means re-running this evaluation by hand every time either lab ships an update; routing each task to whichever model actually wins its category is the alternative. Platforms like Metir AI are built around exactly that model-agnostic approach, giving teams access to Gemini, GPT, Claude and Grok models side by side so a split verdict like this one becomes a routing decision rather than a re-platforming project.

The takeaway

Gemini 3.8 Flash is a genuine improvement over 3.7 Flash on nearly every benchmark Google published, and it beats both Claude Opus 5 and GPT-5.6 Sol outright on several domain-specific agent tasks that matter to real teams. It is also clearly behind Opus 5 on long-horizon computer use, and the token-consumption caveat Google attached to the release means the identical sticker price does not translate into an identical cost per task. Gemini 3.8 Flash Cyber is a smaller story on paper, a benchmark that trails the frontier by half a point, but the production evidence from Chrome, Wiz and Google's own vulnerability research team suggests the gated-access model is finding real defenders willing to put it to work.


Route every task to the right model, automatically

Gemini 3.8 Flash wins some categories and loses others, and that split is exactly the case for not standardizing on a single vendor. With Metir AI you get unified access to Gemini, GPT, Claude, Grok and other leading models in one workspace, so each task goes to whichever model actually performs best on it. Try Metir AI free and let the right model handle every task.

Sources:

  • 3.8 Flash and 3.8 Flash Cyber | Google Blog
  • Gemini API pricing | Google AI for Developers
  • Gemini 3.8 Flash Benchmarks | officechai

Image credits

Header image: Google's Gradient Canopy office building on its Mountain View, California campus, photographed by GualdimG via Wikimedia Commons, licensed under CC BY-SA 4.0. In-body photograph, a wider exterior view of the same building, also by GualdimG via Wikimedia Commons, licensed under CC BY-SA 4.0. Neither photo depicts the Gemini 3.8 Flash or 3.8 Flash Cyber launch itself; both show Google's real Mountain View office campus.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational