Soul-AILab/SoulX-FlashTalk

SoulX-FlashTalk is the first 14B model to achieve sub-second start-up latency (0.87s) while maintaining a real-time throughput of 32 FPS on an 8xH800 node.

What it solves

SoulX-FlashTalk enables the real-time, infinite streaming of audio-driven avatars. It addresses the challenge of creating high-quality talking head animations from audio input that can run continuously without degradation in performance or quality.

How it works

The project utilizes a technique called "Self-Correcting Bidirectional Distillation" to achieve its streaming capabilities. It is built upon base models such as InfiniteTalk and Wan, and incorporates distillation techniques from DMD and Self-forcing++ to optimize the generation process for real-time performance.

Who it’s for

This tool is designed for developers and researchers working on digital humans, virtual avatars, and real-time interactive video systems.

Highlights

  • Infinite Streaming: Capable of generating continuous audio-driven avatar animations.
  • Real-Time Performance: Optimized for high-speed inference, with specific support for multi-GPU setups (e.g., 8xH800) for maximum speed.
  • Large Model Scale: Features a 14B parameter model for high-fidelity output.
  • Flexible Deployment: Supports single-GPU inference with CPU offloading to reduce VRAM requirements.

Related

  • Project
  • Project
  • Project
  • Project
  • Project