Google slashed the cost of running AI agents on July 21 with a new generation of Gemini models that use fewer tokens while maintaining or improving output quality across coding, research, and security tasks.
The numbers tell a clear story. On DeepSWE, a benchmark for real-world software engineering, the new model resolves 49 percent of issues compared to 37 percent from the prior generation. Machine learning research benchmarks climbed from 49.7 percent to 63.9 percent. Early adopters including design platform Figma and legal AI firm Harvey reported smoother document parsing and code migration after switching.
Developers who need maximum throughput on a budget can tap 3.5 Flash-Lite, which runs at 350 output tokens per second at $0.30 per million input tokens. Google designed it for high-volume agentic search and document processing pipelines. A third variant, 3.5 Flash Cyber, teams a security-focused model with the CodeMender agent to find and fix software vulnerabilities automatically.
Safety upgrades include enhanced Frontier Safeguards that make the models substantially more resistant to jailbreaks targeting chemical, biological, and nuclear topics. Google confirmed 3.5 Pro remains in partner evaluation and that work on Gemini 4, described internally as the company’s most ambitious pre-training effort yet, is already running.