bytedance/SALMONN

SALMONN family: A suite of advanced multi-modal LLMs

What it solves

SALMONN is a family of multi-modal large language models designed to give LLMs generic hearing abilities and advanced audio-visual perception. It addresses the limitation of traditional LLMs by enabling them to understand and reason about audio, speech, and video content in a unified way.

How it works

The project provides a suite of specialized models within the same family, each focusing on different multi-modal capabilities:

  • SALMONN: The base model focusing on generic hearing abilities.
  • video-SALMONN (1 & 2): Audio-visual LLMs that can generate high-quality video captions and perform video question answering (QA).
  • ELLSA: An end-to-end model that unifies vision, speech, text, and action in a streaming full-duplex framework for concurrent generation and perception.
  • Speech Quality Assessment: Specialized versions for evaluating speech quality with natural language reasoning.

Who it’s for

Researchers and developers working on multi-modal AI, audio-visual understanding, and end-to-end perception-action systems.

Highlights

  • Multi-modal Integration: Unifies audio, speech, video, and text.
  • Diverse Model Variants: Includes models for general hearing, video captioning, and streaming full-duplex interaction.
  • Comprehensive Data: Provides annotations for 3-stage training, including SQA/AQA data and audio-based storytelling.
  • Open Source: Releases model checkpoints, inference code, and specialized datasets like QualiSpeech.

Related

  • Project
  • Project
  • Project
  • Project
  • Project