Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A New Era for Scalable AI Agents

Google has announced the release of its latest advancements in artificial intelligence with the introduction of three new Gemini models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. These models are engineered to address the growing demand for efficiency, speed, and reliability in building and deploying AI agents at scale. The announcement signifies a strategic push by Google to empower developers and enterprises with more cost-effective and performant AI solutions, particularly for complex agentic workflows.

The new models build upon the foundation of Gemini 3.5 Flash, aiming to optimize token usage, reduce latency, and enhance overall reliability, crucial factors for businesses looking to integrate AI agents into their production environments. This release also comes with a forward-looking statement about the ongoing development of Gemini 3.5 Pro, which is currently in partner testing, and the commencement of pre-training for the next-generation Gemini 4, signaling Google’s sustained commitment to pushing the boundaries of AI capabilities.
Gemini 3.6 Flash: Enhanced Efficiency and Quality

Gemini 3.6 Flash represents a significant leap forward from its predecessor, Gemini 3.5 Flash. This new iteration has been meticulously developed in response to direct feedback from developers and customers, focusing on delivering superior performance in coding and knowledge-based tasks while simultaneously improving token efficiency. Google reports that Gemini 3.6 Flash consumes approximately 17% fewer output tokens compared to 3.5 Flash on the Artificial Analysis Index. This reduction in token consumption translates to fewer reasoning steps and tool calls required to complete multi-step workflows, leading to more streamlined and cost-effective operations.
A key highlight of Gemini 3.6 Flash is its enhanced economic profile. Priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, it offers a lower cost than 3.5 Flash. This pricing strategy is designed to directly reduce the overall cost per agentic task, making the development and deployment of AI agents more accessible and financially viable for a wider range of businesses.

Performance Gains Across Diverse Use Cases
Despite its increased efficiency, Gemini 3.6 Flash demonstrates notable performance improvements across various applications. Early benchmarks and case studies highlight its capabilities in areas such as financial data analysis, code migration, and 3D workflow development. For instance, in financial data analysis and transcript processing, 3.6 Flash, when utilized with Managed Agents on AIS, has shown enhanced efficiency and accuracy compared to 3.5 Flash. Similarly, in code migration tasks orchestrated by AGY, 3.6 Flash exhibits lower latency and higher quality outcomes.

The model’s versatility is further showcased in its application within the Gemini App for developing photographic texture extractors for 3D workflows, and its ability to build interactive theme studios with strong visual understanding skills when paired with AGY and the tldraw offline editor. These examples underscore the model’s adaptability and its potential to revolutionize creative and technical fields.
Commitment to Safety and Security

Google emphasizes that Gemini 3.6 Flash is being released with robust safety enhancements. The model incorporates advanced Frontier Safety safeguards, particularly in domains like Chemical, Biological, Radiological, and Nuclear (CBRN) threats and cyber offense misuses. These safeguards are designed to make the model significantly more resistant to adversarial attacks and jailbreaking attempts, while simultaneously minimizing instances of unwarranted refusals for legitimate and beneficial uses. This focus on safety is crucial for the responsible deployment of powerful AI technologies.
For detailed technical specifications and performance data, Google has made the Gemini 3.6 Flash model card publicly available, offering transparency and further insights into its capabilities.

Gemini 3.5 Flash-Lite: Scaling Agentic Workflows
Complementing the release of 3.6 Flash, Google has also introduced Gemini 3.5 Flash-Lite. This model is specifically optimized for low-latency tasks and scenarios where high throughput is paramount, such as agentic search and extensive document processing. It is positioned as the fastest model within the Gemini 3.5 series, achieving a remarkable speed of 350 output tokens per second, according to Artificial Analysis benchmarks.

Priced at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens, Gemini 3.5 Flash-Lite offers a compelling price-to-performance ratio. This makes it an attractive option for developers and businesses managing high volumes of production traffic. Its quality is reported to be significantly better than the previous 3.1 Flash-Lite, further solidifying its value proposition.
Enhanced Capabilities and Performance Metrics

Gemini 3.5 Flash-Lite is designed to enable efficient scaling of agentic systems. It demonstrates substantial improvements over 3.1 Flash-Lite across various "thinking levels," allowing developers to configure the model to prioritize either low-latency, low-cost execution for high-volume tasks or to engage higher thinking levels for processing more complex, multi-step subagent workloads. The integration of computer use as a built-in tool enhances its reliability in supporting agentic tasks across different platforms.
Performance metrics indicate significant gains in areas such as coding and agentic tasks, with a notable improvement on the Terminal-Bench 2.1 benchmark (54% vs. 31%). It also excels in long-context understanding, achieving 72.2% on the GDM-MRCR v2 benchmark compared to 60.1% with its predecessor, and shows enhanced real-world task execution with a score of 1140 versus 642 on GDPval-AA v2. In fact, on several agentic and coding evaluations, 3.5 Flash-Lite even surpasses the performance of the standard 3 Flash model, including on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%). This positions it as a faster and more capable alternative for workloads previously handled by 2.5 and 3 Flash.

Early adopters of Gemini 3.5 Flash-Lite have lauded its synergy of speed, intelligence, and cost-effectiveness for scaling agentic workflows and data processing. Testimonials from companies like Ashler, Palo Alto Networks, and Ramp highlight its transformative impact on their operations.
Gemini 3.5 Flash Cyber: Securing Software at Scale

The third new model, Gemini 3.5 Flash Cyber, is specifically fine-tuned for cybersecurity applications. Recognizing that the speed of AI in identifying vulnerabilities often outpaces the speed of human remediation, this model is built to tackle the growing threat of software insecurity. Developed on the foundation of 3.5 Flash, Gemini 3.5 Flash Cyber is optimized for detecting, validating, and patching code security issues with exceptional efficiency and at a lower price per token than larger, more general-purpose models.
Within Google’s CodeMender platform, which leverages multiple 3.5 Flash Cyber agents working collaboratively to produce consolidated security reports, the model achieves competitive performance on the widely recognized CyberGym benchmark. This demonstrates its capability in identifying and addressing sophisticated cybersecurity vulnerabilities.

Responsible Deployment of Cybersecurity AI
Given the dual-use nature of cybersecurity AI, Google has adopted a deliberate and cautious approach to the deployment of Gemini 3.5 Flash Cyber. The model will be exclusively accessible to government entities and trusted partners through a limited-access pilot program via CodeMender. This strategy aims to equip frontline defenders with advanced tools to proactively identify and fix critical vulnerabilities, thereby mitigating their potential exploitation, while simultaneously establishing robust controls against broader misuse.

Availability and Future Outlook
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available starting immediately, empowering developers to begin integrating these advanced AI capabilities into their applications and workflows. Google has expressed its eagerness to receive feedback from early users to further refine future iterations of the Gemini models.

Looking ahead, the company has indicated that Gemini 3.5 Pro will be made broadly available once it reaches optimal readiness, and the ambitious pre-training for Gemini 4 has already begun. This continuous development pipeline underscores Google’s long-term vision for AI advancement and its commitment to delivering cutting-edge solutions to its global user base. The ongoing evolution of the Gemini family of models signals a significant trajectory towards more intelligent, efficient, and secure AI applications across a multitude of industries.







