NVIDIA/flashdreams

high-performance inference and serving library for interactive autoregressive video and world models

What it solves

FlashDreams provides a high-performance inference and serving library designed for interactive autoregressive video and world models. It addresses the need for real-time generation in complex applications like autonomous vehicle simulation, robotics, gaming, and virtual environments where low-latency, high-throughput video synthesis is required.

How it works

It serves as an optimized runtime platform that supports a variety of autoregressive models. It provides a unified system for launching models in different modes (such as local-window or WebRTC for interactive demos) and includes first-party integrations for several video and world-model families, including Wan2.1, OmniDreams, and Cosmos-Predict2.5.

Who it’s for

This tool is built for developers and researchers working on real-time world-model applications, specifically those with access to high-end hardware (NVIDIA GPUs with 80 GB VRAM or more, such as the H100).

Highlights

  • Broad Model Support: Integrates multiple families including T2V (text-to-video), I2V (image-to-video), and HDMap-conditioned driving models.
  • Interactive Capabilities: Supports real-time interactive driving demos via WebRTC or local windows.
  • High-Performance Stack: Optimized for CUDA 13.x and PyTorch 2.11.0+.
  • Extensible: Includes developer guides for adding new methods and integrating new models into the pipeline.

Related

  • Project
  • Project
  • Project
  • Project
  • Project