boundless-large-model/boundless-world-model

High-fidelity world models for general embodied intelligence, such as data engines and world simulators.

What it solves

BWM provides a low-cost, high-fidelity simulator for robotic manipulation. It addresses the challenge of predicting physically consistent future video frames based on specific robot actions, which is essential for training and testing robot learning algorithms without needing constant real-world interaction.

How it works

Built upon the Wan2.2-TI2V-5B video diffusion backbone, BWM operates as an action-conditioned world model. It autoregressively predicts future observation chunks by taking an initial observation, dynamic history, and an action chunk as inputs. Robot actions are integrated into the diffusion process via cross-attention tokens and action-conditioned timestep embeddings.

Who it’s for

This project is designed for researchers and developers in robotics and embodied AI who need a physically consistent video simulator to model robot-object interactions, test control strategies, and evaluate robot learning performance.

Highlights

  • Physically Consistent: Maintains long-horizon physical consistency and stable contacts during complex tasks like stacking or hinge interactions.
  • Action-Conditioned: Generates videos based on specific robot action sequences, allowing for precise simulation of manipulation tasks.
  • High Performance: Ranked 1st among open-source models on Track 1 and Track 2 Data Engine of the CVPR 2026 WorldArena Challenge.
  • Generalization: Capable of rolling out future frames even when starting from novel initial scenes or object appearance shifts.

Related

  • Project
  • Project
  • Project
  • Project
  • Project