Gemini Omni Flash Release
Gemini Omni Flash Release
Google DeepMind has announced Gemini Omni Flash, a natively multimodal model that enables the creation and conversational editing of high-quality videos. The model integrates Gemini's reasoning capabilities with creative generation, allowing users to produce video content grounded in real-world knowledge and physical laws.
Conversational Video Editing
Gemini Omni Flash allows users to edit videos using natural language instructions that build upon previous turns. The model maintains consistency in characters, physics, and scene continuity across multiple edits.
- Environmental Transformation: Users can change specific elements or the entire world within a video (e.g., transforming a sculpture into bubbles).
- Action Reimaging: The model can modify existing actions, add new characters or objects, or transform specific moments within a recorded video.
- Iterative Refinement: Users can adjust the environment, camera angle, style, or specific details through a multi-turn conversation without losing the original scene's context.
Knowledge-Grounded Video Generation
Gemini Omni Flash leverages Gemini's internal knowledge of history, science, and cultural context to move beyond simple pattern matching toward meaningful storytelling.
- Improved Physics: The model demonstrates an intuitive understanding of kinetic energy, gravity, and fluid dynamics to create more realistic motion, such as a marble rolling on a chain reaction track.
- Complex Visual Explainers: Omni can generate educational content from short prompts, such as a claymation-style stop-motion video explaining protein folding.
- Contextual Creativity: The model can blend language and imagery to follow complex instructions, such as creating a rapid-fire sequence of 26 items representing the alphabet with specific visual markers and accompanying music.
Multimodal Input Integration
Gemini Omni Flash can synthesize a single cohesive video output from any combination of the following inputs:
- Text: Natural language descriptions of the scene and style.
- Images: Reference images of characters, scenes, or drawings to ensure visual alignment with a specific vision.
- Video: Existing video clips used as a starting point or style reference.
- Audio: Voice references are supported at launch, with other audio input types planned for future release.
Digital Avatars and Safety
To enable personalized content creation, users can utilize the Avatars feature to generate videos that look and sound like them using their own voice.
To ensure responsible AI development, Google has implemented the following safeguards:
- SynthID Watermarking: All videos generated by Gemini Omni include an imperceptible digital watermark to verify their AI origin.
- Verification Tools: Users can verify AI-generated content via the Gemini app, Google Search, and Gemini in Chrome.
- Controlled Audio Editing: The capability to edit audio and speech within videos is currently undergoing further testing to ensure responsible deployment.
Availability
Gemini Omni Flash is rolling out to the following platforms:
- Subscribers: Available to Google AI Plus, Pro, and Ultra subscribers globally via the Gemini app and Google Flow.
- YouTube Users: Rolling out at no cost to users on YouTube Shorts and the YouTube Create App starting the week of May 17, 2026.
- Developers: API access for developers and enterprise customers will be available in the coming weeks.
Sources
- OriginalIntroducing Gemini Omni