AMAP-ML/DreamX-World

DreamX-World: A General-Purpose Interactive World Model

What it solves

DreamX-World addresses the challenge of creating high-fidelity, interactive world simulations that users can explore and modify in real-time. It solves common issues in long-term video generation such as identity, background, and style drift, while enabling precise control over camera movement and environmental changes.

How it works

The model is trained using a scalable data engine that combines Unreal Engine data, gameplay footage, and real-world videos with camera estimation. It employs a progressive training pipeline: first learning fine-grained action control, then open-ended event responses, and finally using Reinforcement Learning to enhance visual fidelity and interaction consistency. To maintain scene persistence when revisiting areas, it uses geometry-guided memory retrieval to recover visual evidence from previous observations. For efficiency, the model utilizes forcing and distillation to make interactive generation practical.

Who it’s for

This project is designed for researchers and developers working on world models, interactive AI simulations, and high-fidelity video generation for applications in gaming, virtual environments, and robotics simulation.

Highlights

  • Interactive Exploration: Supports both first-person and coherent third-person views with stable camera-follow behavior.
  • Long-Horizon Generation: Capable of generating coherent videos up to 1 minute long without significant visual drift.
  • Promptable Events: Allows users to trigger single or compositional events that dynamically transform the environment.
  • Scene Persistence: Uses geometry-guided memory to ensure that revisited regions maintain their layout and object identities.
  • Diverse Environments: Generates a wide range of settings, from realistic urban and natural scenes to stylized sci-fi and fantasy worlds.

Related

  • Project
  • Project
  • Project
  • Project
  • Project