WeChatCV/Stand-In

[CVPR2026 🎉] Stand-In is a lightweight, plug-and-play framework for identity-preserving video generation.

What it solves

Stand-In addresses the challenge of maintaining consistent identity (face and subject similarity) in AI-generated videos. It provides a way to ensure that a specific person or subject from a reference image remains recognizable and consistent across the generated video frames without requiring the massive computational overhead of training a full video model from scratch.

How it works

It is a lightweight, plug-and-play framework that integrates with base text-to-video (T2V) models (such as Wan2.1 and Wan2.2). Instead of retraining the entire model, Stand-In only trains approximately 1% of the additional parameters. This allows it to act as an identity control layer that can be combined with other tools like LoRAs for stylization or VACE for pose and depth control.

Who it’s for

This tool is designed for creators and researchers working with video generation who need high-fidelity subject consistency, including those performing tasks like face swapping, subject-driven video generation, or stylized video production.

Highlights

  • Extreme Efficiency: Only requires training 1% of the base model's parameters.
  • Versatile Integration: Works with community LoRAs for stylization and VACE for pose-guided generation.
  • Multi-Subject Support: Capable of preserving identities for both humans and non-human subjects.
  • Broad Application: Supports text-to-video generation, video face swapping, and identity-preserving stylization.

Related

  • Project
  • Project
  • Project
  • Project
  • Project