|

Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens

Google has launched Gemini 3.7 Flash, the latest mannequin in its Flash tier, three weeks after Gemini 3.6 Flash. The model card describes it as a refinement of three.6 Flash with algorithmic enhancements to the core reasoning basis — not a brand new pretraining run. It accepts textual content, photographs, audio, and video throughout a 1M-token context window, returns as much as 64K output tokens, and helps customizable considering configurations that commerce high quality in opposition to value and latency. The information cutoff stays at March 2026. The features focus in three locations: software program engineering, document-heavy information work, and internet improvement. The sharper argument is value. Gemini 3.7 Flash ships at $0.75 per 1M enter tokens and $3.75 per 1M output tokens — half the unique 3.6 Flash listing fee, and roughly a 3rd the blended value of Claude Sonnet 5 or GPT-5.6 Terra.

Is it Deployable?

Yes, API and enterprise solely. There aren’t any open weights. Access runs by means of hosted surfaces: the Gemini API and Google AI Studio, Google Antigravity, Android Studio, the Gemini Enterprise Agent Platform, and the Gemini Enterprise app. Consumers attain it by means of Gemini Spark on Google AI Pro and Ultra plans.

  • Company match: Startups and mid-market groups acquire probably the most, as a result of the introductory value makes always-on brokers inexpensive and not using a Pro-tier funds. Regulated enterprises get a ruled path by means of Gemini Enterprise. Teams with data-residency or air-gap necessities are excluded — there may be nothing to self-host.
  • Industries: Google’s personal eval set factors at authorized, monetary companies, biosciences, and enterprise operations. The Harvey LAB-AA, GDP.pdf, and AutomationBench outcomes are the tells.
  • Applications: Long-running coding brokers, document-heavy back-office automation, UI technology from screenshots or design methods, and PDF-to-structured-data pipelines.

The Benchmark Picture

On FrontierCode 1.1 Main, which measures manufacturing code high quality, Gemini 3.7 Flash scores 43.6% in opposition to 34.4% for 3.6 Flash. On DeepSWE v1.1, a long-horizon software program engineering eval, it reaches 65.3%. On WebDev Arena it posts an Elo of 1588 versus 1538, the highest rating in Google’s comparability desk.

Document and workflow outcomes transfer additional. GDP.pdf, an skilled PDF comprehension eval, goes from 22.0% to 34.0%. AutomationBench, a non-public enterprise workflow set, goes from 17.0% to 30.4% — forward of each Claude Sonnet 5 at 10.7% and GPT-5.6 Terra at 23.6%. Long-context retrieval on GDM-MRCR v2 at 128k reaches 97.0%.

GPT-5.6 Terra is forward on DeepSWE (69.6%), Terminal-bench 2.1 (87.4%), Terminal-bench 3.0 (20.8%), and OSWorld-2.0 (50.2%). On GDPval-AA v2 information work, 3.7 Flash scores 1525 Elo in opposition to 1598 for Sonnet 5 and 1628 for Muse Spark 1.2. CharXiv Reasoning is a regression: 84.5% with out instruments, down from 85.2% for 3.6 Flash. On the Artificial Analysis Intelligence Index, 3.7 Flash scores 56, in opposition to 57 for each GPT-5.6 Terra and Muse Spark 1.2.