Gemini 3.7 Flash makes a big jump on coding and agents
Google released Gemini 3.7 Flash three weeks after 3.6 Flash, and the jump is larger than the version number suggests. On DeepSWE v1.1, a long-horizon software engineering benchmark, it scores 65.3% against 49.0% for 3.6 Flash. Production code quality on FrontierCode 1.1 rises from 34.4% to 43.6%, and a business workflow benchmark nearly doubles, from 17.0% to 30.4%. Google describes it as its "most intelligent workhorse model yet for coding and agents," and the gains cluster in exactly those areas: multi-step planning, fewer retries, and more functional web layouts (its WebDev Arena Elo moves from 1538 to 1588).
Pricing runs the other way. Through the end of 2026 the model costs $0.75 per million input tokens and $3.75 per million output, half the rate of 3.6 Flash. From January 2027 that doubles to $1.50 and $7.50. Google says it added safeguards against CBRN and cyber misuse while keeping normal use open. The announcement does not name a context window or latency figure, so the "workhorse" speed claim is harder to pin down from the numbers alone.
Why it matters
If you build agents or coding tools on a Flash-tier budget, the DeepSWE and automation scores are the ones to run against your own tasks before trusting the headline, and the introductory price gives you a few months to test and switch before the rate doubles in January.