Articles

Gemini 3.7 Flash: Price, Benchmarks, and What's New

Google shipped Gemini 3.7 Flash with big coding and agent gains, at an introductory $0.75 per million input tokens — half the old Flash price. The details.

Chisato Chisato · · 5 min read
Glowing purple neural fibers converging, representing a large language model

Google’s fast-model cadence is accelerating. On August 13, 2026, the company released Gemini 3.7 Flash, its latest workhorse model, roughly three weeks after the previous Flash update and stable from day one — shipped under the model ID gemini-3.7-flash with no preview suffix. The pitch is familiar but sharpened: a cheaper, faster model that closes much of the gap to frontier systems on coding and agentic tasks, aimed at developers who run models at volume.

The release lands while Google’s flagship Gemini 3.5 Pro remains delayed, making Flash — not Pro — the tip of Google’s model lineup for the moment. It also arrives days after OpenAI previewed its Ultrafast tier for GPT-5.6 Sol, underscoring how the competitive front has shifted from raw capability to the economics of serving models at scale.

What’s new in 3.7 Flash

Gemini 3.7 Flash is positioned as a coding and agent model first. Google says it delivers substantial improvements across software engineering, web development, and agentic workflows over its predecessor, Gemini 3.6 Flash. The specs are squarely in modern-Flash territory:

  • Multimodal input: text, images, audio, and video.
  • Context window: 1 million tokens in, up to 64K tokens out.
  • Configurable thinking: developers can dial reasoning effort up or down, trading answer quality against cost and latency.
  • Full tooling at launch: function calling, structured output, code execution, context caching, the batch API, and search grounding.

The configurable-thinking control is the practical headline for teams building agents. It lets a single model serve both cheap, high-throughput calls and slower, higher-quality reasoning passes without swapping model families — the kind of knob that matters when you are orchestrating thousands of steps.

The benchmarks

Google leaned on coding and agent evaluations to make its case, and the jumps over 3.6 Flash are large:

  • DeepSWE v1.1, a software-engineering benchmark, rose from 49.0% to 65.3%.
  • FrontierCode 1.1 (Main) improved from 34.4% to 43.6%.
  • On web development, Google says the model “generates more functional layouts and feature-complete apps in fewer prompts,” posting an Elo of 1588 on Arena.ai’s WebDev Arena.

A 16-point gain on a software-engineering suite from one Flash generation to the next is unusually steep, and it reflects where the frontier labs are concentrating effort: real-world coding and multi-step agent tasks, the workloads enterprises are actually paying for. As always, vendor-reported benchmarks warrant independent replication — but the direction is consistent with the broader push toward agentic coding models across the industry.

The price is the story

The number developers will fixate on is cost. Through December 31, 2026, Gemini 3.7 Flash is available at an introductory price of $0.75 per million input tokens and $3.75 per million output tokens — roughly half what the previous Flash model cost at its own launch. Starting January 1, 2027, the rate rises to $1.50 and $7.50 per million tokens.

That introductory pricing is a deliberate weapon. Halving the entry price on a model that also posts double-digit benchmark gains compresses the value proposition of rival “workhorse” tiers, and it fits a market that has spent 2026 in a sustained price war among the frontier labs. For the highest-volume workloads — code assistants, document processing, batch classification, agent loops — token price often decides which model gets deployed, and Google is betting a lower floor pulls that traffic to Gemini.

Flash as the flagship, for now

There is a strategic wrinkle in shipping a strong Flash model while the flagship stalls. Gemini 3.5 Pro has slipped past its expected window, and Google has been pre-training the next major generation in the background. Releasing 3.7 Flash keeps Google visibly ahead on cadence and price even without a new Pro at the top of the stack — and for a large share of production traffic, Flash-class models are what teams run anyway. The reasoning-heavy Pro tier is reserved for the hardest problems; the volume lives in Flash.

The speed-and-cost framing also mirrors where competitors are pushing. OpenAI’s Ultrafast tier for GPT-5.6 Sol, which runs the full model far faster on specialized hardware, is a bet that latency is a product feature. Google’s answer with 3.7 Flash is different in kind — cheaper tokens and configurable thinking rather than raw speed — but it targets the same customer: the developer deciding, at scale, which model is worth the per-call cost.

What it means

Gemini 3.7 Flash is less a leap in raw intelligence than a move on the economics of running AI in production, and that is exactly where the 2026 competition is being fought. By pairing a large coding-benchmark gain with an introductory price at half the old Flash rate, Google is trying to make the default, high-volume choice for developers a Google model.

Who wins. Developers and startups running models at scale are the clear beneficiaries: cheaper tokens and configurable thinking make it easier to build agents and coding tools without a frontier-tier bill. Google wins if the lower price pulls sustained API traffic onto its stack, deepening the developer lock-in that comes from context caching, tooling, and search grounding.

Who feels pressure. Rival workhorse tiers — OpenAI’s mid-range GPT models, Anthropic’s faster Claude tiers, and open-weight challengers competing on cost — now face a well-benchmarked model at an aggressive introductory price. Every such cut squeezes the margin on inference for everyone selling it.

What to watch. First, whether independent evaluations reproduce the coding and agent gains Google reported — vendor benchmarks are a starting point, not a verdict. Second, the January 1 price step-up: the introductory rate doubles at year-end, so the real competitive question is what Flash costs in 2027, not just this quarter. Third, the still-delayed Gemini 3.5 Pro — how long Google can lead on Flash cadence while its flagship waits is the strategic subplot behind an otherwise routine-looking release. The near-term signal is unambiguous: the model war has moved decisively onto price and throughput, and Google just pushed the floor lower.

Chisato Chisato · · 6 min read

Gemini 3.6 Flash: Price, Benchmarks, and What's New

Google shipped three new Gemini models—3.6 Flash, 3.5 Flash-Lite, and a security-tuned 3.5 Flash Cyber—while its flagship 3.5 Pro slips and Gemini 4 pre-training begins.

#AI #Google #LLM