shengshu-ai/Vidu-S
Vidu S: Real-Time Interactive, Editable, and Spatial Video Generation
What it solves
Vidu S is a suite of real-time interactive video generation models designed to overcome the latency and stability issues associated with high-resolution video streaming. It enables the creation of high-fidelity digital characters, live video editing, and immersive spatial video that can respond to user instructions in real-time.
How it works
The project consists of two primary versions:
- Vidu S1: A model for voice-controlled digital characters that generates 540p video at up to 42 FPS, supporting custom characters and voice tones.
- Vidu S2: An advanced version that extends capabilities to 720p interactive avatars (25–42 FPS) and real-time editing of incoming video streams.
Vidu S2 utilizes a technique called Self-Replay Forcing (SRF) to prevent error accumulation across streaming segments, ensuring long-horizon stability. For performance, it uses an optimized serving stack (TurboDiffusion and TurboServe) featuring low-bit GEMM and multi-GPU pipelining to run on low-cost GPUs.
Who it’s for
Developers and creators who need real-time, interactive video generation for digital avatars, VR/AR immersive displays, and live video style transfer or character replacement.
Highlights
- Real-time Interactive Avatars: Generates 720p video at 25-42 FPS with support for large body motions like dancing.
- Live Video Editing: Supports style transfer, virtual try-on, and background/character replacement from text instructions while preserving motion.
- Spatial Video: Converts streams into synchronized left- and right-eye views for VR headsets.
- Infinite-length Generation: Capable of continuous real-time generation without stability degradation.
- Efficient Inference: Optimized for low-cost GPUs via TurboDiffusion and TurboServe.
Related
- Project
- Project