Google DeepMind Veo: Integrating Generative Video into Live-Action Filmmaking for ANCESTRA

Google DeepMind has partnered with director Eliza McNitt and Primordial Soup—a storytelling innovation venture founded by Darren Aronofsky—to produce the short film "ANCESTRA." The project demonstrates the integration of Veo, Google's state-of-the-art video generation model, with traditional live-action filmmaking to create complex, high-fidelity sequences that would otherwise be prohibitively expensive or technically difficult to capture.

AI-Driven Pre-Production and Concept Art

To establish the visual foundation of "ANCESTRA," the production team used a combination of Gemini, Imagen, and Veo to translate personal history into cinematic imagery.

  • Gemini for Prompt Engineering: The team uploaded personal photographs of McNitt's birth to Gemini, which described the images in precise aesthetic detail. These descriptions served as the primary prompts for subsequent image and video generation.
  • Imagen for Concept Art: Imagen was used to generate key concept art that defined the film's overall mood, style, and color palette.
  • Veo for Animation: The concept art generated by Imagen was then animated using Veo, with additional text prompts used to guide specific actions and movements.

Advanced Veo Capabilities for Cinematic Control

To meet the requirements of professional filmmaking, Google DeepMind developed new capabilities for Veo to allow for greater personalization and precise control over camera motion.

Personalized Video Generation

To maintain consistent art direction across scenes—specifically for footage of a baby in utero—the team fine-tuned an Imagen model to match specific reference images. Gemini was then used to refine prompts for realistic imagery, which Veo converted into animated scenes via its image-to-video capability. This process ensured that the AI-generated subject remained visually consistent throughout the film.

Motion Matched Video Generation

Veo was utilized to execute precise camera movements that would be difficult to achieve with text prompts alone:

  • Virtual Camera Tracking: For a sequence traveling through the human body into the womb, the team created a 3D model of a human body and recorded a draft shot with a virtual camera. Veo tracked this draft shot's motion to generate the final video.
  • Reference Motion Matching: To depict organic holes closing, the team provided Veo with reference videos of that specific motion. Veo combined these reference motions with text prompts to generate new, high-quality scenes in minutes, bypassing the complex and time-consuming nature of traditional CGI.

Blending Generative Video with Live-Action Footage

To avoid the "uncanny valley" often associated with VFX babies, the production team combined live-action performances with generative elements:

  • Object Addition: Using Veo's "add object" capability, the team inserted a generated newborn baby into live-action footage. By providing Veo with the original footage, a text prompt, and a defined area for the baby, the model generated the subject while keeping the rest of the scene consistent.
  • VFX Integration: These generated elements were then refined using traditional VFX and color grading to ensure a seamless blend between the AI-generated baby and the live-action environment.

Integration with Traditional VFX Pipelines

Generative AI was not used as a replacement for traditional workflows but as a complementary tool. For complex shots—such as a point-of-view shot from inside a hatching crocodile egg at sunset—the team generated multiple individual images and videos using Veo and Imagen, then composed them using traditional VFX compositing techniques.

This partnership between Google DeepMind and Primordial Soup is the first of three planned films, aimed at ensuring generative AI tools are developed based on the actual needs and workflows of professional filmmakers.

Sources