Gemini 3.8 Flash Benchmarks: What Google Published
The one score Google gave, the claims it made without numbers, and why you should still run your own pilot
Google published one benchmark score when it launched Gemini 3.8 Flash: HLE-Verified at 54.9%. It described the model's other gains without providing numbers (Google). Google announced the model on September 2, 2026 (Google).
This page separates what Google actually measured from what it only claimed. We will not invent a score that Google did not publish.
The final sections explain how to interpret vendor benchmarks and test the model on your own workloads instead.
Gemini 3.8 Flash Benchmarks: What Google Reported
Google reported one numeric benchmark for Gemini 3.8 Flash and framed the rest as leadership claims without scores (Google). The table below shows what Google published for each test.
Read the table as a launch-day snapshot from the vendor (Google). It shows the direction Google claims. It does not tell you how the model behaves on your tasks, and most rows carry no number to check.
| Benchmark | What Google reported | What it measures |
|---|---|---|
| HLE-Verified | 54.9% (Google) | Multi-step reasoning |
| DeepSWE v1.1 | Leads most larger frontier models, no score (Google) | End-to-end software fixes |
| Vals Finance Agent V2 | Beats 3.7 Flash and frontier models, no score (Google) | Finance agent tasks |
| Harvey's Legal Agent Benchmark | Beats frontier models, no score (Google) | Legal agent tasks |
| Gray Swan (prompt injection) | Large robustness gain, no score (Google) | Resistance to prompt injection |
Run Your AI On Mac Studio

The ultimate machine for running AI models on your own desk: M5 Max, a 32-core GPU, and 36GB of unified memory.
The One Score: HLE-Verified at 54.9%
HLE-Verified is the only benchmark Google gave a number for, and Gemini 3.8 Flash scored 54.9% (Google). Higher is better on this test.
The benchmark measures multi-step reasoning, the kind of problem that needs several linked steps to reach a correct answer. Google leans on this result because it positions the model around reasoning and coding (Google).
A score of 54.9% means the model got a little over half of these hard reasoning tasks right. That is a signal of strength, not a guarantee, since the model still misses close to half.
The base rate is the part buyers skip. A test built to be hard can leave a strong model failing four tasks in ten, so 54.9% is a real result and a reminder to keep a human check on high-stakes reasoning.
Google's announcement did not publish the full methodology for this test. So confirm the exact task set and scoring on Google's own page before you quote the number, and check whether the same test was run on the models you want to compare.
The Claims Google Made Without Numbers
Google stated three more results as leadership claims with no score attached (Google). Each names a real benchmark, but you cannot check a number that was not published.
On DeepSWE v1.1, a test of fixing real software issues end to end, Google says Gemini 3.8 Flash outperforms most larger frontier models (Google). No percentage was given.
On Vals AI's Finance Agent V2 and Harvey's Legal Agent Benchmark, Google says the model beats 3.7 Flash and other frontier models (Google). Again, no scores were shown.
On prompt-injection resistance, measured with Gray Swan's tests, Google reports a large robustness gain (Google). Better resistance matters most for agents that read untrusted input, since it lowers the odds a poisoned page hijacks the agent. Treat every one of these as a claim to verify, not a settled result.
- DeepSWE v1.1: leads most larger frontier models, no score (Google)
- Vals Finance Agent V2: beats 3.7 Flash and frontier models, no score (Google)
- Harvey's Legal Agent Benchmark: beats frontier models, no score (Google)
- Gray Swan prompt injection: a large robustness gain, no score (Google)
Important: These Are Google's Own Results
Every result on this page is Google's own, not an independent test (Google). That is not a knock on Google. It is how almost all launch-day benchmarks work across the industry.
No third party has published a same-generation head-to-head yet. There is no independent comparison of Gemini 3.8 Flash against rival models, so you cannot rank it against GPT or Claude from these results alone.
Vendors also pick which benchmarks to show, and which numbers to attach. A model can lead on the tests a company highlights and trail on tests it left out.
When we ship model-launch page families across our portfolio at Layer3 Labs, buyers ask first about real-task fit, then read the launch chart. The benchmarks open the conversation. Your own pilot closes it.
How to Read Vendor Benchmarks
Read any vendor benchmark by asking who ran it, what it measured, and whether it gave a number. Those three questions catch most of the traps. Apply them to Gemini 3.8 Flash and every rival model the same way.
First, check the source. If the vendor ran the test, treat it as a claim to verify, not a settled fact. Independent results carry more weight than launch-day charts.
Second, watch for claims with no score. A statement that a model beats frontier models means little without a number and a shared test set (Google). Weigh a claim like that far below a published, reproducible score.
Third, match the benchmark to your job. A high reasoning or legal-agent result means little if your team writes marketing copy. Pick the one or two tests closest to your real work and ignore the rest.
Fourth, ask whether the result is reproducible. A number you can rerun on a public test set is worth more than a claim that names a benchmark but shows no score. When a vendor states a lead without a number, treat it as marketing until an independent run confirms it.
- Ask who ran the test: vendor or independent lab
- Distrust a leadership claim that carries no number
- Confirm what the benchmark actually measures
- Match the test to your real tasks, not the headline
- Prefer a reproducible score over a named-but-unscored claim
Turning One Benchmark Into a Buying Decision
One published score is only a starting point. HLE-Verified at 54.9% tells you how the model handled a broad, expert-level exam (Google), but it says little about the two tasks most teams run on a Flash tier: everyday coding and multi-step agent work. Google described gains there without releasing matched numbers, so the strongest claims for this model are the ones you cannot yet check.
That gap changes how you should read the launch chart. Treat the single HLE-Verified figure as evidence the model is capable in general, then weight it lightly against your own workload, because a broad exam and your codebase are different tests. A model can score well on a reasoning exam and still miss on your prompts, or do the reverse.
The practical move is to bridge the one public number to a private one. Pick 20 to 50 real tasks you already know the right answer to, run them through Gemini 3.8 Flash and through the model you use today, and score both the same way. That private benchmark, on your data, outranks any vendor chart for your decision.
- An exam score shows general capability but not fit for your tasks
- Google published no matched coding or agent numbers for 3.8 Flash
- Build a private benchmark of 20 to 50 known-answer tasks to decide
Run a Short Pilot Before You Switch
The fastest way to trust a benchmark is to replace it with your own test. Pull ten to twenty real tasks your team does every week. Run them through Gemini 3.8 Flash and your current model, then compare the output.
Score the results on what you actually care about: correctness, speed, cost, and how much editing each answer needs. Keep the prompts and the graders the same for both models. That gives you a fair, apples-to-apples read.
This small pilot beats any launch chart for your decision. It uses your data, your workflows, and your quality bar. Google's one score tells you the model is worth testing (Google). The pilot tells you whether to switch.
How to use Gemini 3.8 Flash
You do not host Gemini 3.8 Flash yourself — you use it through a tool, so "getting started" really means choosing the right one.
The fastest way to put Gemini 3.8 Flash to work day to day is inside an AI IDE, and Cursor is the most popular — it supports it directly, so you can be working in minutes. The maker's own option is Antigravity for Gemini 3.8 Flash, if you want the native experience. Prefer a different editor? Windsurf, Zed, and GitHub Copilot drive these models too.
Frequently Asked Questions
- Google published one score, HLE-Verified at 54.9%, a reasoning test (Google). It stated its other results as claims without numbers, saying the model leads on DeepSWE v1.1 and beats rivals on the Vals finance and Harvey legal agent benchmarks. Those carry no published score.
- No. Every result is Google's own from its launch materials (Google). No third party has released a same-generation head-to-head, and there is no independent comparison against rival models. Run your own pilot before you rely on these claims.
- Google says Gemini 3.8 Flash improves on 3.7 Flash across software engineering, agent tasks, and reasoning (Google). It published one number, HLE-Verified at 54.9%, and stated the other gains without scores. The size of the gap is not something the published results settle.
- HLE-Verified measures multi-step reasoning, problems that take several linked steps to solve correctly. Gemini 3.8 Flash scored 54.9% (Google). Google did not publish the full methodology, so verify the task set on Google's page before you quote the number.
- Not reliably. Google's results compare Gemini 3.8 Flash to 3.7 Flash and to unnamed frontier models (Google). No independent, same-generation head-to-head against GPT or Claude has been published. To compare models, run all of them on the same set of your own tasks.
- Run a short pilot. Take ten to twenty real tasks your team does each week, run them through Gemini 3.8 Flash and your current model, and score correctness, speed, cost, and editing effort. Keep the prompts identical for both. This beats any launch claim for your decision.
Want a Real Read on Gemini 3.8 Flash for Your Work?
Book a free 30-minute AI workflow audit with Layer3 Labs. We will help you build a small, fair pilot that tests Gemini 3.8 Flash on your own tasks, so you decide on evidence instead of a launch claim.
Book an Audit