Gemini 3.7 Flash vs Gemini 3 Pro
A task-routing guide to picking the right Google model for each job.
These are both Google models, so this is not an upgrade question. It is a routing question: which one should handle each task? Gemini 3.7 Flash is Google's fast, cheap workhorse for coding and agents. Gemini 3 Pro is Google's Pro-tier model, still in preview as of this writing.
Flash is also the cheaper of the two. During its intro window it runs $0.75 per million input tokens and $3.75 per million output tokens. Gemini 3 Pro costs $2 input and $12 output per million tokens, with $0.20 cached input.
The goal is not to run everything on the most capable model. It is to send high-volume, cost-sensitive work to Flash and save Pro for the harder reasoning. This page gives you a task-routing table to do that.
Flash was released on August 13, 2026. Gemini 3 Pro remains in preview because Google's higher-end Pro flagship was delayed (Google) (press). So the choice today is between a released workhorse and a preview Pro tier, not between two finished models.
Gemini 3.7 Flash vs. Gemini 3 Pro: Side-by-Side
| Dimension | Gemini 3.7 Flash | Gemini 3 Pro |
|---|---|---|
| Best for | High-volume coding, agents, and cost-sensitive work | Harder reasoning where a Pro-tier model is better positioned |
| Availability | Released August 13, 2026 | Preview status as of this writing |
| Input price (per M tokens) | $0.75 intro, then $1.50 from Jan 1, 2027 | $2 |
| Output price (per M tokens) | $3.75 intro, then $7.50 from Jan 1, 2027 | $12 |
| Cached input (per M tokens) | Confirm current caching rate with Google | $0.20 |
| Positioning | "Most intelligent workhorse model yet for coding and agents" | Current-generation Pro-tier model, in preview |
| Deploy today? | Yes, released and available via Spark and the API | Preview status; confirm production-readiness with Google first |
Gemini 3.7 Flash vs Gemini 3 Pro: which should I use?
Use Gemini 3.7 Flash by default, and reach for Gemini 3 Pro only when a task needs heavier reasoning. Both are Google models, so you can route work between them inside the same stack.
Flash is the cheaper option and it is already released. Gemini 3 Pro is in preview as of this writing, which is a reason to test it before you route real traffic to it.
Think of Flash as the model that runs your volume and Pro as the model you call for the tough cases. The rest of this page shows where each line falls.
- Default to Gemini 3.7 Flash for most coding, agent, and drafting work.
- Reach for Gemini 3 Pro on harder reasoning tasks.
- Gemini 3 Pro is in preview, so pilot it before production traffic.
Deciding how to split work between Gemini 3.7 Flash and Gemini 3 Pro? We can design a routing setup that fits your volume, budget, and quality bar.
Book a ConsultationA task-routing table for Google's two models
Match the task to the model, not the model to your pride. High-volume and cost-sensitive jobs go to Gemini 3.7 Flash. Harder reasoning goes to Gemini 3 Pro when Flash is not enough.
The table below is a starting point, not a rule. Your own pilot results should override it, because the right split depends on your tasks and your quality bar.
The pattern holds across most stacks. A small share of requests are genuinely hard, and the rest are routine. Route the routine share to Flash and you cut cost without cutting quality where it matters.
- Bulk code generation and refactors: Gemini 3.7 Flash.
- Agent loops and tool-calling at scale: Gemini 3.7 Flash.
- Web development and debugging at volume: Gemini 3.7 Flash.
- Cost-sensitive, high-request-count features: Gemini 3.7 Flash.
- Deep multi-step reasoning on a hard problem: try Gemini 3 Pro.
- Cases where Flash output is not good enough: escalate to Gemini 3 Pro.
Price: Flash is cheaper, in preview or not
Gemini 3.7 Flash costs less than Gemini 3 Pro on every token. During its intro window Flash runs $0.75 input and $3.75 output per million tokens. Gemini 3 Pro costs $2 input and $12 output per million tokens.
The Flash intro rate holds through December 31, 2026, then rises to $1.50 input and $7.50 output on January 1, 2027. Even at that regular rate, Flash stays cheaper per token than Gemini 3 Pro's published price.
Gemini 3 Pro does list a low $0.20 cached-input rate, which can help on repeated context. Confirm Flash's current caching rate with Google before you model your costs. Prices change without notice, so verify both on Google's pricing page.
The price gap grows with volume. On a large batch of output tokens, Flash's intro rate is roughly a third of Gemini 3 Pro's. That is why routing bulk work to Flash is where most of the savings comes from.
- Gemini 3.7 Flash: $0.75 / $3.75 intro, then $1.50 / $7.50 per million tokens.
- Gemini 3 Pro: $2 input / $12 output per million tokens, $0.20 cached input.
- Flash is cheaper per token in both its intro and regular windows.
When Gemini 3.7 Flash is enough
Gemini 3.7 Flash is enough for most high-volume coding and agent work. Google positions it as its most intelligent workhorse model yet for coding and agents. That is the exact profile of a default model.
Google says Flash gained in software engineering, knowledge work, and web development. It is described as better at debugging and at producing deployable, production-ready code on the first try.
So route bulk code generation, refactors, agent loops, and web-dev tasks to Flash first. It is cheaper and it is built for throughput. Escalate only when a specific task proves too hard.
Google's benchmarks back this fit, though they are Google's own. Flash posts strong coding and web-dev scores against its predecessor. No independent same-generation head-to-head is published, so pilot Flash on your own tasks before you trust the framing.
- High request counts where cost per call matters.
- Coding, debugging, and web development at scale.
- Agent and tool-calling workloads that run all day.
- Knowledge-work drafting where speed and price beat marginal quality.
When to reach for Gemini 3 Pro
Reach for Gemini 3 Pro when a task needs heavier reasoning than Flash delivers. Gemini 3 Pro is Google's current Pro-tier model, positioned for capable, cost-efficient work above the Flash line.
The honest catch is that Gemini 3 Pro is in preview as of this writing. Google's higher-end Pro flagship was delayed, which is part of why Flash arrived first (press). So Pro is the top current option, not the top planned one.
Treat Pro as an escalation target, not a default. Test it on your hard cases, confirm its preview terms with Google, and only then route production traffic to it.
Pro does carry one clear price edge on repeated context. Its cached input runs $0.20 per million tokens. If a hard task reuses the same large context often, that rate can soften the cost of choosing Pro.
- Deep, multi-step reasoning that Flash handles poorly.
- Tasks where a Pro-tier model is better positioned than a workhorse.
- Cases worth a higher per-token price for better output quality.
Why not run everything on the most capable model?
Running every task on the pricier model wastes money on work the cheaper model handles fine. Most requests in a real workload are routine, and Gemini 3.7 Flash is built to serve them cheaply.
The math is simple. Flash's intro output rate is $3.75 per million tokens against Gemini 3 Pro's $12. At high volume, that gap turns into a large bill for quality you may not need on routine calls.
A routing setup captures the savings. Send the bulk to Flash, send the hard minority to Pro, and you pay Pro prices only where they earn their keep.
There is a quality cost to the reverse mistake too. Route a hard reasoning task to Flash and you may pay in poor output, retries, and cleanup. The point of routing is to match each task to the model that fits it best.
- Most calls are routine and do not need a Pro-tier model.
- Flash's output rate is a fraction of Gemini 3 Pro's.
- Routing pays the top price only on the tasks that need it.
Benchmarks: what Google published, and what it did not
Google published Flash-to-Flash benchmark gains, but not a Flash-versus-Pro head-to-head. Its charts compare Gemini 3.7 Flash to Gemini 3.6 Flash, its predecessor. Those numbers are Google's own, not an independent test.
For example, Google reports Gemini 3.7 Flash at 43.6% on FrontierCode 1.1 versus 34.4% for 3.6 Flash, and 65.3% on DeepSWE v1.1 versus 49.0%. It also cites a 1588 WebDev Arena Elo for 3.7 Flash against 1538 for 3.6 Flash.
No independent, same-generation head-to-head between Flash and Gemini 3 Pro has been published. Do not read Flash's gains over its predecessor as a win over Pro. Pilot both on your own tasks before you decide.
- FrontierCode 1.1: Gemini 3.7 Flash 43.6% vs 3.6 Flash 34.4%.
- DeepSWE v1.1: Gemini 3.7 Flash 65.3% vs 3.6 Flash 49.0%.
- WebDev Arena Elo: Gemini 3.7 Flash 1588 vs 3.6 Flash 1538.
- No independent Flash-vs-Pro benchmark exists, so run your own pilot.
How to set up your routing
Default new work to Gemini 3.7 Flash, then escalate the tasks it cannot handle to Gemini 3 Pro. This keeps your bill low and reserves the pricier model for real need.
Run a short pilot before you lock the split. Test a representative sample of your tasks on both models, and confirm Gemini 3 Pro's current preview status and terms with Google first.
Revisit the split as Google ships updates. Pro is in preview and the higher-end flagship was delayed, so the right routing today may change once those land (Google) (press).
Access differs between the two, so plan for it. The Gemini app reaches Flash through Gemini Spark, which needs a Google AI Pro or Ultra plan in 160-plus countries. Developers can reach both through the Gemini API in Google AI Studio.
- Default to Gemini 3.7 Flash; escalate hard tasks to Gemini 3 Pro.
- Pilot both on your own workload before committing.
- Recheck the split as Google updates its Pro line.
The Verdict
For most workloads, Gemini 3.7 Flash should be your default and Gemini 3 Pro your escalation target. Flash is released, cheaper on every token, and built as Google's workhorse for coding and agents. Gemini 3 Pro is the Pro-tier option for harder reasoning, but it is in preview as of this writing.
Do not run everything on the pricier model. Route high-volume, cost-sensitive work to Flash, send the hard minority to Pro, and confirm Pro's preview terms with Google before production traffic. Pilot both on your own tasks, since no independent same-generation head-to-head has been published.
Revisit the split as Google ships updates to its Pro line. The right routing today may shift once the delayed flagship lands. For now, Flash carries the volume and Pro handles the hard cases you send it.
Researched from primary Google documentation and public regulator sources. Pricing and availability are accurate as of Aug 14, 2026 and can change — confirm current terms with each vendor before you buy.
Frequently Asked Questions
- Use Gemini 3.7 Flash by default and reach for Gemini 3 Pro only on harder reasoning tasks. Flash is cheaper and released; Gemini 3 Pro is a Pro-tier model still in preview as of this writing.
- Gemini 3.7 Flash is cheaper on every token. It runs $0.75 / $3.75 per million tokens during its intro window and $1.50 / $7.50 after, versus Gemini 3 Pro's $2 / $12.
- No. Gemini 3 Pro is in preview status as of this writing. Confirm its current availability and production-readiness with Google before routing real traffic to it.
- Google's higher-end Pro flagship was delayed, which is part of why Flash arrived first (press). Gemini 3 Pro remains the current Pro-tier option, in preview as of this writing.
- No, and you should not. Most tasks are routine and Gemini 3.7 Flash handles them cheaply. Route the hard minority to Gemini 3 Pro so you pay its higher price only where it helps.
- No independent, same-generation head-to-head has been published. Google's benchmarks compare Flash to its predecessor, not to Gemini 3 Pro. Pilot both on your own tasks before you decide.
Not sure how to split work between Gemini 3.7 Flash and Gemini 3 Pro?
Book a free 30-minute AI workflow audit with Layer3 Labs. We design a task-routing setup that sends your volume to Gemini 3.7 Flash and reserves Gemini 3 Pro for the hard cases, so you pay top prices only where they earn it.
Book Your Free Audit