UWMILab/UniWM

Official implementation of paper "Unified World Models: Memory-Augmented Planning and Foresight for Visual Navigation"

What it solves

UniWM addresses the challenge of aligning visual foresight (imagining what will happen) with action planning in visual navigation. It replaces modular frameworks with a unified approach to ensure that the robot's decisions are directly grounded in its visually imagined outcomes.

How it works

UniWM uses a single multimodal autoregressive backbone to integrate visual foresight and planning. It employs a hierarchical memory mechanism that combines short-term perceptual cues with long-term trajectory context, allowing the model to maintain stable and coherent reasoning over extended navigation horizons.

Who it’s for

This project is designed for researchers and developers working on embodied AI, autonomous navigation, and world models for robotics.

Highlights

  • Unified multimodal autoregressive backbone for integrated foresight and planning.
  • Hierarchical memory mechanism for balancing short-term and long-term context.
  • Grounded action decisions based on visually imagined outcomes.
  • Includes a dataset hosted on Hugging Face featuring multiple navigation splits (e.g., go_stanford, scand, sacson, recon).

Related

  • Project
  • Project
  • Project
  • Dispatch
  • Project