lartpang/SAMs-CDConcepts-Eval

Inspiring the Next Generation of Segment Anything Models: Comprehensively Evaluate SAM and SAM 2 with Diverse Prompts Towards Context-Dependent Concepts under Different Scenes

What it solves

This project provides a comprehensive evaluation framework to test how well Segment Anything Models (SAM and SAM 2) handle "context-dependent" (CD) concepts. While these models excel at segmenting common objects like cars or people, they often struggle with concepts that rely on surrounding context, such as medical lesions, industrial defects, or camouflaged objects.

How it works

The framework evaluates SAM and SAM 2 across 11 different CD concepts using 2D and 3D images and videos. It employs three prompting strategies—manual, automatic, and intermediate self-prompting—to test the models' discriminative capabilities. Additionally, it includes prompt robustness testing to simulate real-world scenarios where prompts might be imperfect.

Who it’s for

This tool is designed for computer vision researchers and developers working on image and video segmentation, specifically those looking to understand the performance boundaries of foundation models in specialized domains like medicine and industry.

Highlights

  • Evaluates both SAM and SAM 2 across natural, medical, and industrial scenes.
  • Supports multiple modalities, including 2D/3D images and videos.
  • Implements a unified framework for manual, automatic, and intermediate self-prompting.
  • Includes robustness testing for imperfect prompts to simulate real-world usage.

Related

  • Project
  • Project
  • Project
  • Project