AMAP-ML/DreamX-Creator
Democratizing Native Audio-Video Generation at 2K Resolution
What it solves
DreamX-Creator addresses the challenge of generating high-resolution video and synchronized audio simultaneously. It solves the problem of misalignment between visual and auditory streams and the lack of native 2K resolution in joint generation frameworks.
How it works
The system uses a base generator that jointly models modality-specialized video and audio streams. It employs Gated Cross-Modal Attention and Progressive Joint Training to allow bidirectional interaction between audio and video. To improve quality, it integrates Audio-Video Reinforcement Learning with Modality-Aware Multimodal Feedback for better semantic consistency and synchronization. Finally, an Autoregressive 1-Step 2K Refiner upgrades the output to 2K resolution while maintaining the original content and timing.
Who it’s for
This framework is designed for researchers and developers working on native multimodal generation, specifically those needing synchronized audio-video content at high resolutions.
Highlights
- Native Joint Generation: Simultaneously models audio and video streams rather than generating them sequentially.
- Bidirectional Interaction: Uses Gated Cross-Modal Attention to ensure audio-video synchronization.
- High Resolution: Includes a dedicated 1-step refiner to upscale videos to 2K resolution.
- RL-Enhanced: Utilizes reinforcement learning and multimodal feedback to refine visual and audio quality.
Related
- Project
- Project
- Project
- Project