OpenGVLab/InternVideo
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
What it solves
InternVideo provides a series of foundation models designed to improve how machines understand and reason about video content. It addresses the challenge of scaling multimodal understanding, enabling AI to process long-horizon video data and perform complex contextual reasoning.
How it works
The project evolves through a series of iterations:
- InternVideo (v1): Uses generative and discriminative learning for general video foundation models.
- InternVideo2: Scales these models for better multimodal video understanding.
- InternVideo2.5: Focuses on long and rich context modeling for video multimodal large language models (MLLMs).
- InternVideo3: Implements efficient long-horizon agents for multimodal contextual reasoning.
- InternVideo-Next: Aims for genuine world understanding through general video foundation models.
It is supported by InternVid, a large-scale video-text dataset used for both understanding and generation.
Who it’s for
This is primarily for AI researchers and engineers working on video understanding, multimodal learning, and the development of video-centric AI agents.
Highlights
- Iterative Model Series: A progression from general foundation models to specialized long-context and agentic reasoning models.
- Large-Scale Data: Includes InternVid, a dataset containing up to 230 million video-text pairs.
- Diverse Model Sizes: Offers various model scales, including distilled smaller versions (S/B/L) for efficiency.
- Agent Integration: Recent versions incorporate long-horizon agents for complex reasoning tasks.
Related
- Project
- Project
- Project
- Project
- Project