Gemini 3.7 Flash: Price, Benchmarks, and What's New
Google shipped Gemini 3.7 Flash with big coding and agent gains, at an introductory $0.75 per million input tokens — half the old Flash price. The details.
Google’s fast-model cadence is accelerating. On August 13, 2026, the company released Gemini 3.7 Flash, its latest workhorse model, roughly three weeks after the previous Flash update and stable from day one — shipped under the model ID gemini-3.7-flash with no preview suffix. The pitch is familiar but sharpened: a cheaper, faster model that closes much of the gap to frontier systems on coding and agentic tasks, aimed at developers who run models at volume.
The release lands while Google’s flagship Gemini 3.5 Pro remains delayed, making Flash — not Pro — the tip of Google’s model lineup for the moment. It also arrives days after OpenAI previewed its Ultrafast tier for GPT-5.6 Sol, underscoring how the competitive front has shifted from raw capability to the economics of serving models at scale.
What’s new in 3.7 Flash
Gemini 3.7 Flash is positioned as a coding and agent model first. Google says it delivers substantial improvements across software engineering, web development, and agentic workflows over its predecessor, Gemini 3.6 Flash. The specs are squarely in modern-Flash territory:
- Multimodal input: text, images, audio, and video.
- Context window: 1 million tokens in, up to 64K tokens out.
- Configurable thinking: developers can dial reasoning effort up or down, trading answer quality against cost and latency.
- Full tooling at launch: function calling, structured output, code execution, context caching, the batch API, and search grounding.
The configurable-thinking control is the practical headline for teams building agents. It lets a single model serve both cheap, high-throughput calls and slower, higher-quality reasoning passes without swapping model families — the kind of knob that matters when you are orchestrating thousands of steps.
The benchmarks
Google leaned on coding and agent evaluations to make its case, and the jumps over 3.6 Flash are large:
- DeepSWE v1.1, a software-engineering benchmark, rose from 49.0% to 65.3%.
- FrontierCode 1.1 (Main) improved from 34.4% to 43.6%.
- On web development, Google says the model “generates more functional layouts and feature-complete apps in fewer prompts,” posting an Elo of 1588 on Arena.ai’s WebDev Arena.
A 16-point gain on a software-engineering suite from one Flash generation to the next is unusually steep, and it reflects where the frontier labs are concentrating effort: real-world coding and multi-step agent tasks, the workloads enterprises are actually paying for. As always, vendor-reported benchmarks warrant independent replication — but the direction is consistent with the broader push toward agentic coding models across the industry.
The price is the story
The number developers will fixate on is cost. Through December 31, 2026, Gemini 3.7 Flash is available at an introductory price of $0.75 per million input tokens and $3.75 per million output tokens — roughly half what the previous Flash model cost at its own launch. Starting January 1, 2027, the rate rises to $1.50 and $7.50 per million tokens.
That introductory pricing is a deliberate weapon. Halving the entry price on a model that also posts double-digit benchmark gains compresses the value proposition of rival “workhorse” tiers, and it fits a market that has spent 2026 in a sustained price war among the frontier labs. For the highest-volume workloads — code assistants, document processing, batch classification, agent loops — token price often decides which model gets deployed, and Google is betting a lower floor pulls that traffic to Gemini.
Flash as the flagship, for now
There is a strategic wrinkle in shipping a strong Flash model while the flagship stalls. Gemini 3.5 Pro has slipped past its expected window, and Google has been pre-training the next major generation in the background. Releasing 3.7 Flash keeps Google visibly ahead on cadence and price even without a new Pro at the top of the stack — and for a large share of production traffic, Flash-class models are what teams run anyway. The reasoning-heavy Pro tier is reserved for the hardest problems; the volume lives in Flash.
The speed-and-cost framing also mirrors where competitors are pushing. OpenAI’s Ultrafast tier for GPT-5.6 Sol, which runs the full model far faster on specialized hardware, is a bet that latency is a product feature. Google’s answer with 3.7 Flash is different in kind — cheaper tokens and configurable thinking rather than raw speed — but it targets the same customer: the developer deciding, at scale, which model is worth the per-call cost.
What it means
Gemini 3.7 Flash is less a leap in raw intelligence than a move on the economics of running AI in production, and that is exactly where the 2026 competition is being fought. By pairing a large coding-benchmark gain with an introductory price at half the old Flash rate, Google is trying to make the default, high-volume choice for developers a Google model.
Who wins. Developers and startups running models at scale are the clear beneficiaries: cheaper tokens and configurable thinking make it easier to build agents and coding tools without a frontier-tier bill. Google wins if the lower price pulls sustained API traffic onto its stack, deepening the developer lock-in that comes from context caching, tooling, and search grounding.
Who feels pressure. Rival workhorse tiers — OpenAI’s mid-range GPT models, Anthropic’s faster Claude tiers, and open-weight challengers competing on cost — now face a well-benchmarked model at an aggressive introductory price. Every such cut squeezes the margin on inference for everyone selling it.
What to watch. First, whether independent evaluations reproduce the coding and agent gains Google reported — vendor benchmarks are a starting point, not a verdict. Second, the January 1 price step-up: the introductory rate doubles at year-end, so the real competitive question is what Flash costs in 2027, not just this quarter. Third, the still-delayed Gemini 3.5 Pro — how long Google can lead on Flash cadence while its flagship waits is the strategic subplot behind an otherwise routine-looking release. The near-term signal is unambiguous: the model war has moved decisively onto price and throughput, and Google just pushed the floor lower.
Keep reading
Chisato · · 6 min read Gemini 3.8 Flash: Benchmarks, Price, and a Cyber Sibling
Google shipped Gemini 3.8 Flash, its third Flash release in six weeks, holding the $0.75 input price and adding a locked-down 3.8 Flash Cyber variant.
Chisato · · 6 min read Gemini 3.6 Flash: Price, Benchmarks, and What's New
Google shipped three new Gemini models—3.6 Flash, 3.5 Flash-Lite, and a security-tuned 3.5 Flash Cyber—while its flagship 3.5 Pro slips and Gemini 4 pre-training begins.
Chisato · · 6 min read Google DeepMind WeatherNext 3: Hourly 5km AI Forecasts
Google DeepMind's WeatherNext 3 delivers hourly, up-to-5km AI weather forecasts with 15-day, 64-member ensembles and gains in cyclone prediction.