Soul-AILab/SoulX-FlashTalk
SoulX-FlashTalk is the first 14B model to achieve sub-second start-up latency (0.87s) while maintaining a real-time throughput of 32 FPS on an 8xH800 node.
What it solves
SoulX-FlashTalk enables the real-time, infinite streaming of audio-driven avatars. It addresses the challenge of creating high-quality talking head animations from audio input that can run continuously without degradation in performance or quality.
How it works
The project utilizes a technique called "Self-Correcting Bidirectional Distillation" to achieve its streaming capabilities. It is built upon base models such as InfiniteTalk and Wan, and incorporates distillation techniques from DMD and Self-forcing++ to optimize the generation process for real-time performance.
Who it’s for
This tool is designed for developers and researchers working on digital humans, virtual avatars, and real-time interactive video systems.
Highlights
- Infinite Streaming: Capable of generating continuous audio-driven avatar animations.
- Real-Time Performance: Optimized for high-speed inference, with specific support for multi-GPU setups (e.g., 8xH800) for maximum speed.
- Large Model Scale: Features a 14B parameter model for high-fidelity output.
- Flexible Deployment: Supports single-GPU inference with CPU offloading to reduce VRAM requirements.
Related
- Project
- Project
- Project
- Project
- Project