Introduction of Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Our latest Gemini models deliver the efficiency, latency, and reliability needed to build AI agents at scale. Developers and customers require higher token efficiency, lower latency, and more reliable performance. Our Flash series models are designed to strike the sweet spot of efficiency and quality to enable scalable agentic workflows.
3.6 Flash: The Workhorse Model
3.6 Flash builds on 3.5 Flash, delivering improved coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash, with reductions observed up to 65% in benchmarks like DeepSWE by Datacurve, all at a lower cost per output token.
// Sample code: Performance comparison of 3.6 Flash
float efficiency3_6 = 1 - (tokens_used_3_6 / tokens_used_3_5);
3.5 Flash-Lite: The Fastest Model
3.5 Flash-Lite is the fastest model in the 3.5 series, delivering 350 output tokens per second, significantly outperforming prior Flash-Lite generations. Priced at $0.3/1M input tokens and $2.5/1M output tokens, it offers a strong price-to-performance ratio for developers.
3.5 Flash Cyber: Efficient Cybersecurity Solutions
3.5 Flash Cyber is tailored for finding and fixing security vulnerabilities, achieving competitive performance on the popular benchmark CyberGym within CodeMender, which utilizes multiple 3.5 Flash Cyber agents.
Conclusion
Gemini 3.6 Flash and the 3.5 series models are now available, inviting developers to explore through Google AI Studio and the Gemini app. We look forward to the release of 3.5 Pro soon.
Blogger's Review: The launch of Gemini 3.6 Flash and the 3.5 series models represents a significant leap in AI agent technology, particularly in optimizing efficiency and cost, which will greatly enhance developer productivity and the scalability of applications.