Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: Efficiency, Speed, and Security Focus

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: Efficiency, Speed, and Security Focus

Overview of the New Gemini Flash Models

Google introduced three new models in the Gemini Flash series: Gemini 3.6 Flash as a workhorse model for coding, knowledge work, and multimodal tasks; Gemini 3.5 Flash-Lite as the fastest, most cost‑effective 3.5‑class model for high‑throughput agentic workflows; and Gemini 3.5 Flash Cyber paired with the CodeMender agent for cybersecurity vulnerability detection and remediation.

Token Efficiency and Cost Improvements in Gemini 3.6 Flash

Gemini 3.6 Flash reduces output token usage by 17% compared to Gemini 3.5 Flash according to the Artificial Analysis Index, and in some benchmarks such as DeepSWE by Datacurve it achieves up to 65% lower token consumption. The model is priced at $1.50 per million input tokens and $7.50 per million output tokens, a reduction from the $9 per million output tokens of Gemini 3.5 Flash.

Performance gains reported for Gemini 3.6 Flash include:

  • DeepSWE: 49% vs 37% (precision with fewer unwanted code edits)
  • MLE Bench: 63.9% vs 49.7%
  • OSWorld‑Verified (computer use): 83.0% vs 78.4%
  • GDPval‑AA v2: 1421 vs 1349

The model also ships with enhanced Frontier Safety safeguards for CBRN and cyber‑offense misuse, making it more resistant to jailbreaks while minimizing refusals for beneficial uses.

High‑Throughput Performance of Gemini 3.5 Flash‑Lite

Gemini 3.5 Flash‑Lite delivers 350 output tokens per second as measured by the Artificial Analysis Index, making it the fastest model in the 3.5 series. It is priced at $0.3 per million input tokens and $2.5 per million output tokens.

The model enables low‑latency, high‑volume agentic workloads and shows improvements over prior Flash‑Lite generations:

  • Terminal‑Bench 2.1: 54% vs 31%
  • GDM‑MRCR v2: 72.2% vs 60.1%
  • GDPval‑AA v2: 1140 vs 642
  • Outperforms Gemini 3 Flash on SWE‑Bench Pro (54.2% vs 49.6%) and OSWorld‑Verified (74.0% vs 65.1%)

Computer use is available as a built‑in client‑side tool via the Gemini API and Gemini Enterprise.

Specialized Cybersecurity Model: Gemini 3.5 Flash Cyber in CodeMender

Gemini 3.5 Flash Cyber is a fine‑tuned version of Gemini 3.5 Flash focused on finding and fixing cybersecurity vulnerabilities. It is deployed alongside the CodeMender code security agent, which runs multiple agents to produce a combined report. On the CyberGym benchmark the combination reaches competitive performance at the frontier.

Because of the dual‑use nature of the technology, access to Gemini 3.5 Flash Cyber is limited to governments and trusted partners through a pilot program via CodeMender.

Availability and Integration Paths

Developers can start using Gemini 3.6 Flash and Gemini 3.5 Flash‑Lite today through:

  • Gemini API via Google AI Studio and Android Studio
  • Google Antigravity
  • Gemini Enterprise Agent Platform and the Gemini Enterprise app
  • The Gemini app for end‑users

Gemini 3.5 Flash‑Lite is also rolling out in Google Search. Gemini 3.5 Pro is currently in partner testing with a broader release planned, and Google has begun pre‑training for Gemini 4.

Community Reception and Comparative Observations

Hacker News comments reflect both appreciation for the speed and cost improvements and concerns about pricing trends, model comparisons, and productization.

Some users praised the throughput and cost‑effectiveness:

"3.5 Flash lite is pretty comparable to Opus 4.8 (at least for the couple tests I did) while simultaneously being 6x faster and 19x cheaper." – @michaelbuckbee

Others highlighted the model’s speed for iterative work:

"I use antigravity from time to time and in my experience, 3.5 flash is an underrated model, so long as you know what it’s good for. It’s very good at frontend (much better than gpt 5.5) and it’s fast, so it’s a great tool for iteration." – @postalcoder

Pricing increases drew criticism:

"Gemini 2.5 Flash‑Lite has been my go to for cheap document processing at scale (especially with 50% off batch mode), but they are really boiling the frog with pricing increases with each version: gemini‑2.5‑flash‑lite: $0.10 input / $0.40 output gemini‑3.1‑flash‑lite: $0.25 input / $1.50 output gemini‑3.5‑flash‑lite: $0.30 input / $2.50 output (a 6.25x increase over 2.5!) Now watch them deprecate Gemini 2.5 Flash‑Lite in the coming months..." – @JeremyHerrman

Comparisons to other models were frequent:

"It is both less intelligent and more expensive than GLM‑5.2, while being closed weight." – @jgbuddy "GLM 5.2 is better, also cheaper, and almost as fast. So essentially, a big L for Google." – @zwaps "So it’s GLM‑5.2 performance for almost twice the price. That said, the speed looks really good." – @resonious

Users also noted product‑experience issues:

"Google somehow managed to snatch defeat from the jaws of success with their AI products. They literally forced me and my company out of Antigravity by phasing out AI Ultra subscription without any proper product follow‑up." – @stonewhite "Google has not changed. Following two facts are like tautologies by now. 1. Their AI efforts are very fundamental research oriented. They are really good at it. 2. Their productization sucks. You should never build anything build around Google only APIs, AI or not." – @u1hcw9nx

Some comments pointed out missing features or ecosystem gaps:

"Really, what’s up with Gemini still not supporting connectors/MCPs/plugins/whatever‑they’re‑called‑this‑month on web? It makes it a non‑starter for any kind of serious use." – @lilytweed "I have just tried to switch to 3.6 instead of 3.5 in antigravity and it seems to constantly spit ‘critical instruction: STOP CALLING TOOLS NOW. YOU MUST WAIT FOR WAKEUP.’ I think I will switch back to 3.5." – @sega_sai

Overall, the release emphasizes efficiency, throughput, and a specialized security offering, while the community response highlights both the practical benefits of the new Flash models and ongoing concerns about pricing, model competitiveness, and Google’s productization approach.

Sources