Gemini 3.7 Flash Benchmarks: The 5 Scores, Explained
What Google's five published numbers measure, how they beat Gemini 3.6 Flash, and why you should still run your own pilot.
Google published five benchmark scores for Gemini 3.7 Flash on launch day, August 13, 2026 (Google). Each score beats the same test on Gemini 3.6 Flash (Google). This guide lists all five in one table, then explains what each one measures in plain English.
One thing matters before you read on. Every number here comes from Google's own charts (Google). None of them are independent. No outside lab has published a same-generation head-to-head yet.
So treat these results as a strong signal, not proof. The last section shows how to read vendor benchmarks and how to test the model on your own work.
Gemini 3.7 Flash benchmarks: all five scores
Gemini 3.7 Flash posted higher scores than Gemini 3.6 Flash on all five benchmarks Google published (Google). The gains are largest on the coding and agent tests. The table below puts both models side by side.
Read the table as a launch-day snapshot from the vendor (Google). It shows the direction of the upgrade. It does not tell you how the model behaves on your tasks.
| Benchmark | Gemini 3.7 Flash | Gemini 3.6 Flash | What it points at |
|---|---|---|---|
| FrontierCode 1.1 (Main) | 43.6% | 34.4% | Hard coding problems |
| DeepSWE v1.1 | 65.3% | 49.0% | Real software fixes |
| WebDev Arena (Elo) | 1588 | 1538 | Web app building |
| GDP.pdf | 34.0% | 22.0% | Reading long documents |
| AutomationBench | 30.4% | 17.0% | Multi-step agent tasks |
Not sure whether Gemini 3.7 Flash's benchmark gains hold up on your real work? We will help you design a short, fair pilot and read the results straight.
Book a ConsultationThe coding benchmarks: FrontierCode 1.1 and DeepSWE v1.1
The two coding benchmarks test whether the model can write and fix real code, not just answer trivia. Google leans on these because it positions Gemini 3.7 Flash as a coding and agent model (Google). Both scores are pass rates, so higher is better.
FrontierCode 1.1 (Main) measures how often the model solves hard programming problems. The name and Google's framing point at frontier-level coding tasks (Google). Gemini 3.7 Flash scored 43.6%, up from 34.4% on Gemini 3.6 Flash (Google).
DeepSWE v1.1 measures how often the model fixes real software issues end to end. Think of a bug ticket that needs a working patch, not a hint. Gemini 3.7 Flash scored 65.3%, up from 49.0% (Google).
Google's launch blog did not publish the full test methodology for these two benchmarks. So confirm the exact task set and scoring on Google's own page before you quote a number.
The web, document, and agent benchmarks
The other three benchmarks test web building, long-document reading, and multi-step agent work. Together they cover the jobs a business AI usually gets asked to do. Each one improved from Gemini 3.6 Flash to Gemini 3.7 Flash (Google).
WebDev Arena (Elo) rates how well the model builds web apps, scored like a chess rating. A higher Elo means the model wins more head-to-head matchups judged by people. Gemini 3.7 Flash scored 1588, up from 1538 (Google).
GDP.pdf measures how well the model reads and reasons over long documents. The name points at dense, report-style files like PDFs. Gemini 3.7 Flash scored 34.0%, up from 22.0% (Google).
AutomationBench measures how well the model runs multi-step tasks on its own, the kind of work an agent does. Gemini 3.7 Flash scored 30.4%, up from 17.0% (Google). That near-doubling is the biggest relative jump of the five.
As with the coding tests, Google did not publish full methodology for these three in its launch blog. Check Google's page for the current task definitions before you rely on any single figure.
The 3.6 to 3.7 delta: what actually improved
Gemini 3.7 Flash gained the most on the agent and document tests, where scores nearly doubled (Google). This tracks with Google's pitch that the model is better at debugging and at producing deployable code on the first try (Google).
AutomationBench rose from 17.0% to 30.4%, and GDP.pdf rose from 22.0% to 34.0% (Google). Both are large relative jumps from a low base. Low starting scores are easier to move, so read big percentage gains with care.
The coding tests improved too, with DeepSWE v1.1 up 16.3 points and FrontierCode 1.1 up 9.2 points (Google). WebDev Arena moved 50 Elo points (Google), a smaller but real gain on a rating scale.
The pattern is consistent. The model got better at the exact tasks Google is targeting: software fixes, web builds, long documents, and agent workflows (Google). Whether that shows up in your work is a separate question the benchmarks cannot answer.
Important: these are Google's own numbers
Every score on this page is Google's own published result, not an independent test (Google). That is not a knock on Google. It is how almost all launch-day benchmarks work across the industry.
No third party has published a same-generation head-to-head yet. There is no independent SWE-bench Verified comparison of Gemini 3.7 Flash against rival models at this point. So you cannot rank it against GPT or Claude from these numbers alone.
Vendors also pick which benchmarks to show. A model can lead on the five tests a company publishes and trail on tests it left out. That is a reason to test the model yourself, not a reason to distrust Google.
When we ship model-launch page families across our portfolio at Layer3 Labs, the first buyer question is almost always about real-task fit, not the launch chart. Benchmarks open the conversation. Your own pilot closes it.
How to read vendor benchmarks
Read any vendor benchmark by asking who ran it, what it measured, and what it left out. Those three questions catch most of the traps. Apply them to Gemini 3.7 Flash and every rival model the same way.
First, check the source. If the vendor ran the test, treat it as a claim to verify, not a settled fact. Independent results carry more weight than launch-day charts.
Second, match the benchmark to your job. A high coding score means little if your team writes marketing copy. Pick the one or two tests closest to your real work and ignore the rest.
Third, watch the base rate. A jump from 17% to 30% looks huge but still means the model fails most of the time on that task (Google). Small gains near the top of a scale can matter more than large gains near the bottom.
- Ask who ran the test: vendor or independent lab
- Confirm what the benchmark actually measures
- Match the test to your real tasks, not the headline
- Read percentage gains against the starting score
- Note which benchmarks the vendor did not show
Run a short pilot before you switch
The fastest way to trust a benchmark is to replace it with your own test. Pull ten to twenty real tasks your team does every week. Run them through Gemini 3.7 Flash and your current model, then compare the output.
Score the results on what you actually care about: correctness, speed, cost, and how much editing each answer needs. Keep the prompts and the graders the same for both models. That gives you a fair, apples-to-apples read.
This small pilot beats any launch chart for your decision. It uses your data, your workflows, and your quality bar. Google's numbers tell you the model is worth testing (Google). The pilot tells you whether to switch.
Frequently Asked Questions
- Google published five scores for Gemini 3.7 Flash: FrontierCode 1.1 at 43.6%, DeepSWE v1.1 at 65.3%, WebDev Arena at 1588 Elo, GDP.pdf at 34.0%, and AutomationBench at 30.4% (Google). All five beat Gemini 3.6 Flash (Google). These are Google's own numbers, not independent results.
- No. Every published score is Google's own result from its launch materials (Google). No third party has released a same-generation head-to-head yet, and there is no independent SWE-bench Verified comparison against rival models. Run your own pilot before you rely on these numbers.
- Gemini 3.7 Flash improved on all five benchmarks Google published (Google). The biggest jumps were on agent and document tasks: AutomationBench rose from 17.0% to 30.4% and GDP.pdf from 22.0% to 34.0% (Google). DeepSWE v1.1 rose from 49.0% to 65.3% (Google).
- DeepSWE v1.1 measures how often the model fixes real software issues end to end, like resolving a bug ticket with a working patch. Gemini 3.7 Flash scored 65.3%, up from 49.0% on Gemini 3.6 Flash (Google). Google did not publish the full methodology in its launch blog, so verify the task set on Google's page.
- Not reliably. Google's five benchmarks only compare Gemini 3.7 Flash to Gemini 3.6 Flash (Google). No independent, same-generation head-to-head against GPT or Claude has been published. To compare models, run all of them on the same set of your own tasks.
- Run a short pilot. Take ten to twenty real tasks your team does each week, run them through Gemini 3.7 Flash and your current model, and score correctness, speed, cost, and editing effort. Keep the prompts identical for both. This beats any launch chart for your decision.
Want a real read on Gemini 3.7 Flash for your work?
Book a free 30-minute AI workflow audit with Layer3 Labs. We will help you build a small, fair pilot that tests Gemini 3.7 Flash on your own tasks, so you decide on evidence instead of a launch chart.
Book an Audit