nv-tlabs/vipe
ViPE: Video Pose Engine for Geometric 3D Perception
What it solves
ViPE addresses the challenge of extracting precise 3D spatial information from raw, unconstrained videos. It allows users to automatically annotate camera poses and generate dense depth maps without needing specialized hardware, working across various lens types including pinhole, wide-angle, and 360-degree panoramas.
How it works
ViPE acts as a spatial AI engine that estimates three key components from video footage: camera intrinsics (the internal camera settings), camera motion (the path the camera took), and dense near-metric depth maps (the distance of objects from the camera).
Who it’s for
This tool is designed for developers and researchers working in 3D perception, computer vision, and spatial AI who need to transform raw video into structured 3D data.
Highlights
- Broad Lens Support: Works with pinhole, wide-angle, and 360-degree panorama footage.
- Multiple Pipeline Options: Integrates with various pipelines, including Depth-Anything 3 and Lyra.
- High Performance: Recent updates include CUDA fused kernels and pipeline caching for significant speed-ups.
- Open Source: Available via PyPI for easy installation.
Related
- Project
- Project
- Project
- Project
- Project