KangLiao929/Puffin

Puffin Series: Towards Unified Multimodal 3D World Models

What it solves

Puffin addresses the challenge of creating unified multimodal models that can both understand and generate 3D worlds. It moves beyond simple image generation by enabling camera-centric spatial intelligence, allowing the system to understand and generate the world from arbitrary viewpoints and orientations.

How it works

The project consists of two main iterations:

  1. Puffin: A unified multimodal model focused on camera-centric understanding and generation.
  2. Puffin-World: An expanded version that represents worlds through three native 3D world states: physics, geometry, and appearance. This allows the model to perform camera-to-world understanding and image- or text-to-3D world generation without needing external perception or reconstruction modules.

Who it’s for

Researchers and developers working on 3D world modeling, spatial intelligence, and multimodal AI that integrates vision and 3D spatial awareness.

Highlights

  • Unified Framework: Combines understanding and generation in a single model.
  • Native 3D States: Uses physics, geometry, and appearance to represent the world.
  • Camera-Controllable: Supports generation based on specific camera viewpoints and orientations.
  • Large-Scale Datasets: Includes the Puffin-4M and Puffin-16M datasets for training and evaluation.

Related

  • Project
  • Project
  • Project
  • Project
  • Project