Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Release
Google has introduced a new suite of efficient models—Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber—designed to reduce latency and costs for developers building AI agents at scale. The primary focus of this release is improving token efficiency and execution speed to make agentic workflows more economically viable.
Gemini 3.6 Flash: Enhanced Efficiency and Quality
Gemini 3.6 Flash serves as the primary workhorse model, offering improved performance in coding and knowledge work while reducing the number of output tokens required to complete tasks.
Key Performance Improvements
- Token Efficiency: According to the Artificial Analysis Index, 3.6 Flash reduces output token usage by 17% compared to 3.5 Flash. In specific benchmarks like DeepSWE, token reduction is observed up to 65%.
- Coding and Research: The model shows higher precision with fewer unwanted code edits. It achieved 49% on DeepSWE (up from 37% for 3.5 Flash) and 63.9% on MLE Bench (up from 49.7%).
- Computer Use: Computer use is now a built-in client-side tool via the Gemini API and Gemini Enterprise, with OSWorld-Verified scores increasing to 83.0% from 78.4%.
- Knowledge Work: 3.6 Flash outperforms its predecessor in knowledge-based tasks, scoring 1421 vs 1349 on GDPval-AA v2.
Pricing and Safety
Gemini 3.6 Flash is priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. It includes enhanced Frontier Safety safeguards to resist jailbreaks in domains such as Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense, while minimizing refusals for beneficial use cases.
Gemini 3.5 Flash-Lite: High-Throughput Agentic Scaling
Gemini 3.5 Flash-Lite is the fastest model in the 3.5 series, optimized for low-latency tasks and high-volume document processing.
Speed and Performance
- Throughput: The model delivers 350 output tokens per second per the Artificial Analysis Index.
- Agentic Capabilities: 3.5 Flash-Lite outperforms 3.1 Flash-Lite across multiple benchmarks, including Terminal-Bench 2.1 (54% vs 31%) and GDM-MRCR v2 (72.2% vs 60.1%).
- Comparison to 3 Flash: In some agentic and coding evaluations, 3.5 Flash-Lite outperforms 3 Flash, specifically on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%).
Pricing and Configuration
Flash-Lite is priced at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens. Developers can configure the model with different "thinking levels" to balance between low-latency execution for high-volume tasks and higher reasoning for multi-step subagent workloads.
Gemini 3.5 Flash Cyber and CodeMender
Gemini 3.5 Flash Cyber is a specialized model fine-tuned for detecting, validating, and patching cybersecurity vulnerabilities. It is integrated into CodeMender, an agent infrastructure that utilizes multiple 3.5 Flash Cyber agents to generate comprehensive security reports.
Due to the dual-use nature of cybersecurity AI, 3.5 Flash Cyber is not available to the general public. It is restricted to governments and trusted partners via a limited-access pilot program to ensure frontline defenders can fix vulnerabilities before they are exploited.
Future Roadmap
Google has confirmed that Gemini 3.5 Pro is currently testing with partners and will be broadly available soon. Additionally, the company has commenced pre-training for Gemini 4, described as their most ambitious pre-training run to date.
Community Insights and Critical Reception
While the technical specifications highlight efficiency gains, the developer community on Hacker News expressed several concerns regarding Google's product strategy and pricing trends.
Pricing Concerns
Several users noted a trend of increasing costs for "Lite" models over time. One user highlighted the price jump for Flash-Lite versions:
Gemini 2.5 Flash-Lite: $0.10 input / $0.40 output Gemini 3.1 Flash-Lite: $0.25 input / $1.50 output Gemini 3.5 Flash-Lite: $0.30 input / $2.50 output (a 6.25x increase over 2.5!)
Product Ecosystem and Stability
Critics pointed to a fragmented product experience across Vertex AI, AI Studio, Gemini Enterprise, and Antigravity. Some developers reported frustration with abrupt product decisions and the deprecation of older, cheaper models, leading to increased operational costs for price-sensitive workloads.
Competitive Positioning
Some users argued that the models are being outpaced by competitors, specifically mentioning GLM 5.2 as being potentially more intelligent and cheaper, while others praised the sheer speed of the Flash series for iterative frontend development and high-volume agentic tasks.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch