Google launches Gemini 3.7 Flash, halves the cost of its coding workhorse

Google's new Flash model scores higher on coding benchmarks and costs half its standard rate, at least through the end of 2026.

2 min
Google launches Gemini 3.7 Flash, halves the cost of its coding workhorse

Google released Gemini 3.7 Flash on August 13, 2026, its "most intelligent workhorse model yet for coding and agents" and the first major update to the Flash line in three weeks.

Illustration: lower AI inference cost

The headline move is price. Through December 31, 2026 the model costs $0.75 per million input tokens and $3.75 per million output tokens, half of the standard rate that takes effect on January 1, 2027, when those numbers double to $1.50 and $7.50.

The pitch is rare in model launches: better and cheaper at once. Google says 3.7 Flash improves on 3.6 Flash across software engineering, knowledge work, and web development. On its own benchmarks, first-pass code accuracy rose on FrontierCode 1.1 Main from 34.4% to 43.6% and on DeepSWE v1.1 from 49.0% to 65.3%. In web development the model scored an Elo of 1588 on WebDev Arena, up from 1538. For document-heavy work it reached 34.0% on the GDP.pdf benchmark, up from 22.0%, and 30.4% on AutomationBench, up from 17.0%.

Google frames the gains as better adaptation when a task hits a roadblock, clearer intent clarification, and tighter instruction following, the traits that matter most for autonomous coding and business agents. That positions 3.7 Flash against OpenAI's GPT-5.6 Sol, Anthropic's Claude, and a wave of lower-priced open-weight Chinese models in a widening price war.

The discount is temporary, which matters for teams weighing total operating cost. Google argues fewer retries and less manual oversight will offset the higher list price in 2027, but that claim is unproven at scale. The launch also underscores Google's rapid cadence on Flash while its next flagship Pro model stays absent.

Sources: Google, Introducing Gemini 3.7 Flash (Aug 13, 2026) | VentureBeat, Gemini 3.7 Flash 50% price cut (Aug 13, 2026)

Written by

More to read

  • Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude

    Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude Anthropic has published experimental results demonstrating Claude's ability to autonomously design de novo protein binders with physical wet-lab validation and automate complex analytical chemistry workflows. The findings show frontier LLMs acting as autonomous agents across computational biology and molecular characterization pipelines. In the primary experiment, Anthropic evaluated Claude Mythos Previe

    1 min
  • Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture

    Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture Cerebras Systems has announced the CS-4, a rack-scale AI accelerator system designed around three of its next-generation Wafer Scale Engine 3 Turbo (WSE-3 Turbo) chips and a modular hardware architecture dubbed Nexus. Cerebras confirmed that initial customer shipments for the CS-4 are scheduled to begin in the current quarter. The new system marks a structural shift from Cerebras's single-wafer CS

    1 min
  • AI FinOps: Cutting LLM Inference Costs by 30-60% Through Model Tiering, Caching, and GPU Optimization

    AI FinOps: Cutting LLM Inference Costs by 30-60% Through Model Tiering, Caching, and GPU Optimization Inference costs have become the second-largest line item in enterprise AI budgets, trailing only talent spend according to RapidData's State of Enterprise AI 2026. This shift represents a fundamental inversion from the 2021-2023 era when training dominated AI expenditure. The compounding nature of serving costs—accumulating every hour as long as users hit the API—means that even modest producti

    1 min