NoizAI/HelixWorld

🪐 HelixWorld: real-time interactive audio-visual world model.

What it solves

HelixWorld creates a real-time interactive audio-visual world model. It allows users to navigate a scene—walking forward or turning around—where both the visual and audio components update dynamically and synchronously based on the camera's movement and the user's prompts.

How it works

Starting with an image and a text prompt, the model generates a navigable environment. As the user provides movement actions (such as walking forward or turning), the system updates the spatial field of both the picture and the sound, ensuring the audio is not a soundtrack but a spatially aware sound field that follows the viewpoint.

Who it’s for

Developers and researchers interested in world models, interactive AI-generated environments, and synchronized audio-visual generation.

Highlights

  • Joint audio-video generation: Visuals and sound are born together and update in sync.
  • Camera navigation: Supports roaming via actions like W, A, S, D and turning.
  • Spatial sound field: Audio updates dynamically based on the camera's perspective.
  • Interactive demo: Available via a web browser at helixworld.org.
  • Offline inference: Preview v1 provides code and checkpoints for local execution on NVIDIA GPUs.

Related

  • Project
  • Project
  • Project
  • Project
  • Project