shengshu-ai/Vidu-S

Vidu S: Real-Time Interactive, Editable, and Spatial Video Generation

What it solves

Vidu S is a suite of real-time interactive video generation models designed to overcome the latency and stability issues associated with high-resolution video streaming. It enables the creation of high-fidelity digital characters, live video editing, and immersive spatial video that can respond to user instructions in real-time.

How it works

The project consists of two primary versions:

  • Vidu S1: A model for voice-controlled digital characters that generates 540p video at up to 42 FPS, supporting custom characters and voice tones.
  • Vidu S2: An advanced version that extends capabilities to 720p interactive avatars (25–42 FPS) and real-time editing of incoming video streams.

Vidu S2 utilizes a technique called Self-Replay Forcing (SRF) to prevent error accumulation across streaming segments, ensuring long-horizon stability. For performance, it uses an optimized serving stack (TurboDiffusion and TurboServe) featuring low-bit GEMM and multi-GPU pipelining to run on low-cost GPUs.

Who it’s for

Developers and creators who need real-time, interactive video generation for digital avatars, VR/AR immersive displays, and live video style transfer or character replacement.

Highlights

  • Real-time Interactive Avatars: Generates 720p video at 25-42 FPS with support for large body motions like dancing.
  • Live Video Editing: Supports style transfer, virtual try-on, and background/character replacement from text instructions while preserving motion.
  • Spatial Video: Converts streams into synchronized left- and right-eye views for VR headsets.
  • Infinite-length Generation: Capable of continuous real-time generation without stability degradation.
  • Efficient Inference: Optimized for low-cost GPUs via TurboDiffusion and TurboServe.

Related

  • Project
  • Project