bytedance/SALMONN
SALMONN family: A suite of advanced multi-modal LLMs
What it solves
SALMONN is a family of multi-modal large language models designed to give LLMs generic hearing abilities and advanced audio-visual perception. It addresses the limitation of traditional LLMs by enabling them to understand and reason about audio, speech, and video content in a unified way.
How it works
The project provides a suite of specialized models within the same family, each focusing on different multi-modal capabilities:
- SALMONN: The base model focusing on generic hearing abilities.
- video-SALMONN (1 & 2): Audio-visual LLMs that can generate high-quality video captions and perform video question answering (QA).
- ELLSA: An end-to-end model that unifies vision, speech, text, and action in a streaming full-duplex framework for concurrent generation and perception.
- Speech Quality Assessment: Specialized versions for evaluating speech quality with natural language reasoning.
Who it’s for
Researchers and developers working on multi-modal AI, audio-visual understanding, and end-to-end perception-action systems.
Highlights
- Multi-modal Integration: Unifies audio, speech, video, and text.
- Diverse Model Variants: Includes models for general hearing, video captioning, and streaming full-duplex interaction.
- Comprehensive Data: Provides annotations for 3-stage training, including SQA/AQA data and audio-based storytelling.
- Open Source: Releases model checkpoints, inference code, and specialized datasets like QualiSpeech.
Related
- Project
- Project
- Project
- Project
- Project