MCG-NJU/VideoChat3
A generalist video MLLM built for fine-grained motion, long-form reasoning, temporal grounding, and online proactive response.
What it solves
VideoChat3 addresses the challenge of creating a generalist video understanding model that can handle diverse tasks—such as detecting subtle motion, reasoning over hour-long videos, pinpointing precise temporal evidence, and providing proactive responses to live streams—all within a single, efficient architecture.
How it works
It is a 4B parameter multimodal large language model (MLLM) that utilizes an I3D-ViT architecture to achieve 16Ñ spatiotemporal compression, reducing redundant visual tokens. It also employs Adaptive Frame Resolution, which increases resolution only when closer visual inspection is required for specific evidence, allowing for more efficient streaming perception.
Who it’s for
This project is for researchers and developers working on video-based AI, multimodal learning, and efficient video reasoning systems.
Highlights
- Generalist capabilities: A single model capable of motion analysis, long-video reasoning, temporal grounding, and online proactive response.
- Token efficiency: Uses I3D-ViT to compress visual tokens while maintaining essential spatiotemporal evidence.
- Adaptive perception: Dynamically adjusts frame resolution based on the need for detailed visual inspection.
- Fully open: Provides public access to model weights, training code, and complete training datasets.
관련
- 프로젝트
- 프로젝트
- 프로젝트
- 프로젝트
- 프로젝트