Lixsp11/sekai-codebase

[NeurIPS 2025] Sekai: A Video Dataset towards World Exploration

What it solves

Sekai addresses the lack of high-quality, annotated egocentric video data needed for immersive world exploration and generation. It provides a massive scale of first-person and drone-view footage to help AI models better understand real-world continuity, navigation, and the relationship between visuals and audio.

How it works

Sekai is a comprehensive dataset consisting of two primary components: Sekai-Real (sourced from YouTube) and Sekai-Game (sourced from video games). The dataset includes over 5,000 hours of 720p video across 100+ countries and 750+ cities. Each video is paired with detailed annotations, including location, scene, weather, crowd levels, captions, and normalized camera trajectories (intrinsic and extrinsic matrices).

Who it’s for

This project is designed for researchers and developers working on video understanding, autonomous navigation, and video-audio co-generation.

Highlights

  • Massive Scale: Over 5,000 hours of high-resolution video.
  • Global Coverage: Spans more than 100 countries and 750 cities.
  • Diverse Perspectives: Includes both first-person walking and drone perspectives.
  • Rich Annotations: Provides precise camera trajectories and environmental metadata (weather, scene, etc.).
  • Long Sequences: Focuses on sequences 60 seconds or longer to maintain real-world continuity.

Related

  • Project
  • Dispatch
  • Dispatch
  • Dispatch
  • Project