Gemini 2.5 Flash-Lite Release Notes
Google DeepMind has released the stable version of Gemini 2.5 Flash-Lite, the fastest and most cost-efficient model in the Gemini 2.5 family. This model is designed for scaled production use, prioritizing high intelligence per dollar and low latency for tasks such as translation and classification.
Technical Specifications and Pricing
Gemini 2.5 Flash-Lite is the lowest-cost model in the 2.5 series, priced at $0.10 per 1 million input tokens and $0.40 per 1 million output tokens. Additionally, audio input pricing has been reduced by 40% compared to the preview launch.
Key technical capabilities include:
- Context Window: A 1 million-token context window.
- Reasoning: Native reasoning capabilities that can be optionally toggled on for demanding use cases via controllable thinking budgets.
- Tool Integration: Support for native tools including Code Execution, URL Context, and Grounding with Google Search.
- Performance: Lower latency than both 2.0 Flash-Lite and 2.0 Flash across a broad sample of prompts.
Performance Benchmarks and Quality
Gemini 2.5 Flash-Lite demonstrates higher overall quality than 2.0 Flash-Lite across several domains, including:
- Coding
- Mathematics
- Science
- Reasoning
- Multimodal understanding
Production Use Cases
Several organizations have already deployed Gemini 2.5 Flash-Lite to optimize specific operational workflows:
- Satlyt: Utilizes the model for a decentralized space computing platform to handle in-orbit telemetry summarization and satellite-to-satellite communication parsing. This has resulted in a 45% reduction in latency for onboard diagnostics and a 30% decrease in power consumption.
- HeyGen: Employs the model to automate video planning, optimize content, and translate videos into over 180 languages.
- DocsHound: Uses the model to process long videos and extract thousands of screenshots to generate documentation and training data for AI agents.
- Evertune: Leverages the model's speed to scan and synthesize large volumes of model output to provide dynamic insights on brand representation across AI models.
Availability and Migration
Gemini 2.5 Flash-Lite is available via Google AI Studio and Vertex AI. Developers can access the model by specifying gemini-2.5-flash-lite in their code. Users of the preview version should migrate to the stable alias, as the preview alias is scheduled for removal on August 25th.