Google has introduced two new models in its Gemini series: Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, while also providing a glimpse into the forthcoming Gemini 4.
Gemini 3.6 Flash
Building upon the previous release at I/O 2026, Gemini 3.6 Flash incorporates feedback from developers and users to enhance token efficiency across various tasks. Compared to its predecessor, Gemini 3.5 Flash, the new model reduces output token consumption by 17%, as measured by the Artificial Analysis Index. It achieves this by requiring fewer reasoning steps and tool calls for multi-step workflows. Additionally, the pricing has been adjusted to $1.50 per million input tokens and $7.50 per million output tokens, down from the previous $9 per million output tokens.
In coding performance, Gemini 3.6 Flash offers higher precision with fewer unnecessary code edits and reduced execution loops. Notably, it generates more reliable, production-ready code, as evidenced by improvements in DeepSWE (49% compared to 37%) and MLE Bench (63.9% versus 49.7%). For knowledge-based tasks, the model scores 1421 on GDPval-AA, up from 1349. Its computer use capabilities have also improved, with scores increasing from 78.4% on OSWorld-Verified to 83%. Furthermore, the knowledge cutoff date has been updated from January 2025 to March 2026.
Gemini 3.5 Flash-Lite
Designed for high-throughput and low-latency tasks such as agentic search and document processing, Gemini 3.5 Flash-Lite offers significantly better quality than the earlier 3.1 Flash-Lite model released in March. It is priced at $0.30 per million input tokens and $2.50 per million output tokens. The model demonstrates substantial improvements in coding and agentic tasks, with Terminal-Bench 2.1 scores rising from 31% to 54%, long-context understanding as seen in GDM-MRCR v2 increasing from 60.1% to 72.2%, and real-world task execution in GDPval-AA v2 improving from 642 to 1140. Additionally, Gemini 3.5 Flash-Lite outperforms Gemini 3 Flash in SWE-Bench Pro (54.2% versus 49.6%) and OSWorld-Verified (74.0% compared to 65.1%).
Gemini 3.5 Flash Cyber
Google also announced Gemini 3.5 Flash Cyber, a model tailored for identifying and addressing security vulnerabilities. Leveraging the performance and efficiency of the Flash series, this model aims to detect, validate, and patch code security issues at scale, offering a cost-effective solution compared to larger models. Google’s CodeMender tool utilizes multiple 3.5 Flash Cyber agents. Initially, access to this model is limited to governments and trusted partners as part of a pilot program, providing frontline defenders with a proactive approach to mitigating critical vulnerabilities before exploitation.
Availability and Future Developments
Both Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now available in the Gemini app, with the latter also being integrated into Search. Developers can access these models through Google Antigravity, AI Studio, and Android Studio. Looking ahead, Google has indicated that Gemini 3.5 Pro is currently undergoing testing with partners, and further details about Gemini 4 are anticipated in the near future.
The introduction of Gemini 3.6 Flash and 3.5 Flash-Lite underscores Google’s commitment to advancing AI efficiency and accessibility. By reducing token consumption and enhancing performance across various tasks, these models offer developers and users more cost-effective and powerful tools. The forthcoming Gemini 4 is poised to continue this trajectory, potentially setting new standards in AI capabilities.