zai-org/SCAIL-2
Official Implementation of SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning
What it solves
SCAIL-2 addresses the limitations of traditional character animation that rely on intermediate pose representations (like skeleton maps), which often struggle with complex motions, non-human driving sources (e.g., animals), and identity generalization. It aims to provide a more flexible, end-to-end system for character animation and character replacement in videos.
How it works
SCAIL-2 uses a Unified Motion Transfer Interface with specific masking channels and a dedicated RoPE design to bypass intermediate pose representations. It was trained on 60K motion pairs synthesized from other models (SCAIL-Preview, Wan-Animate, and MoCha) using a "reserve driving" approach to develop emergent capabilities. The system supports both animation mode (end-to-end or pose-driven) and replacement mode. It also incorporates Bias-Aware DPO (Direct Preference Optimization) to improve fine details like lip and eye synchronization and reduce hand distortion.
Who it’s for
This project is for developers and creators working with AI-driven video generation, character animation, and video editing who need high-fidelity motion transfer and character replacement capabilities.
Highlights
- End-to-End Animation: Supports character animation without relying on skeleton maps, enabling animal-driving scenarios.
- Character Replacement: Allows replacing a character in a video with a reference image, including a Relighting LoRA to ensure natural lighting and shadow blending.
- Multi-Reference Support: Enables zero-shot multi-reference generation to provide additional visual details (e.g., different views or close-ups) of a character.
- Bias-Aware DPO: Uses a novel mechanism to improve synchronization and detail quality.
- ComfyUI Integration: Available for use within the ComfyUI ecosystem.
Related
- Project
- Project
- Project
- Dispatch
- Dispatch