KangLiao929/Puffin
Puffin Series: Towards Unified Multimodal 3D World Models
What it solves
Puffin addresses the challenge of creating unified multimodal models that can both understand and generate 3D worlds. It moves beyond simple image generation by enabling camera-centric spatial intelligence, allowing the system to understand and generate the world from arbitrary viewpoints and orientations.
How it works
The project consists of two main iterations:
- Puffin: A unified multimodal model focused on camera-centric understanding and generation.
- Puffin-World: An expanded version that represents worlds through three native 3D world states: physics, geometry, and appearance. This allows the model to perform camera-to-world understanding and image- or text-to-3D world generation without needing external perception or reconstruction modules.
Who it’s for
Researchers and developers working on 3D world modeling, spatial intelligence, and multimodal AI that integrates vision and 3D spatial awareness.
Highlights
- Unified Framework: Combines understanding and generation in a single model.
- Native 3D States: Uses physics, geometry, and appearance to represent the world.
- Camera-Controllable: Supports generation based on specific camera viewpoints and orientations.
- Large-Scale Datasets: Includes the Puffin-4M and Puffin-16M datasets for training and evaluation.
Related
- Project
- Project
- Project
- Project
- Project