hujiecpp/PE3R

[CVPR'26] PE3R: Perception-Efficient 3D Reconstruction. Take 2 - 3 photos with your phone, upload them, wait a few minutes, and then start exploring your 3D world via text!

What it solves

PE3R addresses the challenge of creating 3D reconstructions of scenes that are both computationally efficient and semantically aware. It eliminates the need for specialized depth sensors, relying solely on 2D images to build 3D models that can be understood through language.

How it works

The system leverages a combination of advanced vision models to achieve its goals. It integrates MASt3R for 3D reconstruction, SAM (Segment Anything Model) and SAM 2 for segmentation, and SigLIP for language-based semantic understanding, allowing it to perform zero-shot generalization across different scenes and objects.

Who it’s for

This project is designed for researchers and developers working in computer vision, 3D scene reconstruction, and embodied AI who need a fast, efficient way to turn 2D images into semantically labeled 3D environments.

Highlights

  • Input Efficiency: Uses only 2D images for reconstruction.
  • Time Efficiency: Accelerates the process of 3D semantic reconstruction.
  • Zero-Shot Generalization: Works across various scenes and objects without needing scene-specific training.

Related

  • Project
  • Project
  • Project
  • Project
  • Project