Google has introduced new models in its Gemini Flash series aimed at helping developers and businesses build and scale production-grade AI agents with improved token efficiency, lower latency and more reliable performance.
The new lineup includes Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber, which is designed to work with the company’s CodeMender code security agent. The models are targeted at agentic workflows that require AI systems to perform complex tasks while controlling computational costs and response times.
Gemini 3.6 Flash is positioned as a workhorse model for coding, knowledge work and multimodal applications. According to the Artificial Analysis Index, the model uses 17 per cent fewer output tokens than Gemini 3.5 Flash. Google also said that, in certain benchmarks such as DeepSWE by Datacurve, it has observed up to a 65 per cent reduction in output token usage.
The model is also designed to complete multi-step workflows with fewer reasoning steps and tool calls. Priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, Gemini 3.6 Flash is intended to reduce the overall cost of running AI agents.
Google has also launched Gemini 3.5 Flash-Lite, described as its fastest and most cost-effective 3.5-class model. According to the Artificial Analysis Index, the model delivers output speeds of up to 350 tokens per second and is designed to improve performance in agentic workflows compared with earlier Flash-Lite generations.
For cybersecurity applications, Google is introducing Gemini 3.5 Flash Cyber in CodeMender. The offering combines a specialised cybersecurity model with the CodeMender code security agent, with the aim of supporting advanced code security and vulnerability remediation workflows.
The company said Gemini 3.5 Pro is currently being tested with partners and will be made broadly available once ready. Google also confirmed that it has begun its most ambitious pre-training run to date as it works towards the next generation of models, Gemini 4.
The new releases highlight the growing focus among AI model developers on improving the efficiency of AI agents. As businesses increasingly deploy agents for coding, research, cybersecurity and other multi-step tasks, reducing token consumption, latency and operational costs is becoming increasingly important for production-scale AI adoption.


