Gemini Music Generation with Lyria 3
Google DeepMind has integrated the Lyria 3 generative music model into the Gemini app, allowing users to create 30-second AI-generated tracks using text or images. This update expands Gemini's creative toolset beyond images and video to include high-quality, customizable audio generation.
Lyria 3 Capabilities and Features
Lyria 3 is the latest generative music model from Google DeepMind, designed to translate user prompts into catchy, high-quality tracks. It improves upon previous Lyria models in three primary areas: the automatic generation of lyrics based on prompts, increased creative control over tempo, vocals, and style, and the production of more realistic and musically complex tracks.
Users can generate music through two primary input methods:
- Text-to-Track: Users can describe a specific mood, genre, or memory to create instrumental audio or tracks with lyrics. For example, a user can request a "fun afrobeat track" based on a specific personal memory.
- Image and Video-to-Track: Users can upload a photo or video, and Gemini will compose a track with lyrics that fit the mood of the visual content.
Each generated track is 30 seconds long and includes custom cover art created by Nano Banana.
Integration with YouTube Dream Track
Lyria 3 is also being integrated into YouTube's Dream Track. Initially available in the U.S. and now expanding to creators in other countries, the model enhances the quality of soundtracks for YouTube Shorts, allowing creators to generate more customized lyrical verses or backing tracks.
Audio Verification and SynthID
To ensure the identification of AI-generated content, all tracks produced in the Gemini app are embedded with SynthID, an imperceptible watermark. Google has also expanded its verification tools within the Gemini app to include audio. Users can upload an audio file and ask Gemini if it was generated by Google AI; the model will then check for theID watermark and use its own reasoning to provide a response.
Responsible AI Development and Copyright
Google DeepMind states that Lyria 3 is designed for original expression rather than the mimicry of existing artists. When a user prompts the model with a specific artist's name, Gemini treats the request as broad creative inspiration for style or mood rather than a direct imitation.
To protect intellectual property, Google has implemented filters to check outputs against existing content. Users are required to adhere to Google's Terms of Service and Gen AI prohibited use policies, which forbid the violation of intellectual property and privacy rights. A reporting mechanism is also available for content that may violate rights.
Availability and Access
Lyria 3 is available in the Gemini app for users aged 18 and older in English, German, Spanish, French, Hindi, Japanese, Korean, and Portuguese. The feature is rolling out on desktop today and will reach the mobile app over the next several days. Users with Google AI Plus, Pro, and Ultra subscriptions have higher usage limits.