OpenGVLab/InternVideo

[ECCV2024] Video Foundation Models & Data for Multimodal Understanding

What it solves

InternVideo provides a series of foundation models designed to improve how machines understand and reason about video content. It addresses the challenge of scaling multimodal understanding, enabling AI to process long-horizon video data and perform complex contextual reasoning.

How it works

The project evolves through a series of iterations:

  • InternVideo (v1): Uses generative and discriminative learning for general video foundation models.
  • InternVideo2: Scales these models for better multimodal video understanding.
  • InternVideo2.5: Focuses on long and rich context modeling for video multimodal large language models (MLLMs).
  • InternVideo3: Implements efficient long-horizon agents for multimodal contextual reasoning.
  • InternVideo-Next: Aims for genuine world understanding through general video foundation models.

It is supported by InternVid, a large-scale video-text dataset used for both understanding and generation.

Who it’s for

This is primarily for AI researchers and engineers working on video understanding, multimodal learning, and the development of video-centric AI agents.

Highlights

  • Iterative Model Series: A progression from general foundation models to specialized long-context and agentic reasoning models.
  • Large-Scale Data: Includes InternVid, a dataset containing up to 230 million video-text pairs.
  • Diverse Model Sizes: Offers various model scales, including distilled smaller versions (S/B/L) for efficiency.
  • Agent Integration: Recent versions incorporate long-horizon agents for complex reasoning tasks.

Related

  • Project
  • Project
  • Project
  • Project
  • Project