Google introduces Gemini 3.7 Flash, arriving three weeks after 3.6 Flash, positioned as its most capable workhorse model yet for coding and agentic workflows. It shows notable gains over 3.6 Flash on FrontierCode (43.6% vs 34.4%), DeepSWE v1.1 (65.3% vs 49.0%), WebDev Arena Elo (1588 vs 1538), GDP.pdf (34.0% vs 22.0%), and AutomationBench (30.4% vs 17.0%). It launches at an introductory price of $0.75/1M input and $3.75/1M output tokens, roughly half the prior Flash cost, available through end of year. Gemini Spark, the personal agent for AI Pro/Ultra subscribers in over 160 countries, is being upgraded to use 3.7 Flash starting immediately. The model also ships updated safety safeguards for CBRN and cyber misuse domains.

4m read timeFrom blog.google
Post cover image

Questions this post answers

What is Gemini 3.7 Flash and how does it compare to Gemini 3.6 Flash for coding?

Gemini 3.7 Flash is Google's updated workhorse model for coding and agents, released three weeks after 3.6 Flash. It scores 43.6% versus 34.4% on FrontierCode 1.1 Main and 65.3% versus 49.0% on DeepSWE v1.1, showing stronger debugging, issue resolution, and first-pass code accuracy than its predecessor. Developers comparing model versions before switching can track releases like this one on daily.dev.

How much does the Gemini 3.7 Flash API cost per token?

Gemini 3.7 Flash launches at an introductory price of $0.75 per 1 million input tokens and $3.75 per 1 million output tokens, roughly half the cost of the prior 3.6 Flash pricing. This introductory rate is available through the end of the year. Teams budgeting for LLM API costs can follow pricing shifts like this on daily.dev.

How does Gemini 3.7 Flash perform on web development and UI generation tasks?

Gemini 3.7 Flash generates more functional layouts and feature-complete apps in fewer prompts than 3.6 Flash, with strong design adherence when given a reference screenshot, image, or design system. It scores an Elo of 1588 versus 1538 for 3.6 Flash on Arena.ai's WebDev Arena leaderboard. Developers picking a model for UI generation workflows can weigh benchmarks like these on daily.dev.

10 Impressions