visual-layer/fastdup

fastdup is a powerful, free tool designed to rapidly generate valuable insights from image and video datasets. It helps enhance the quality of both images and labels, while significantly reducing data operation costs, all with unmatched scalability.

What it solves

fastdup addresses the challenge of managing and cleaning large-scale visual datasets. It helps users identify and remove duplicates, near-duplicates, outliers, mislabeled images, and low-quality images (such as those that are blurry, too dark, or too bright) to ensure high-quality training data for AI models.

How it works

The tool uses an optimized C++ engine to analyze image and video datasets, whether they are labeled or unlabeled. It can process data locally or in the cloud, supporting the use of custom embeddings from sources like TIMM (PyTorch Image Models) or ONNX models (e.g., DINOv2) to surface dataset issues. It also provides built-in visualization tools to generate galleries of duplicates, outliers, and image statistics.

Who it’s for

It is designed for data scientists and ML engineers who need to curate massive visual datasets—ranging from hundreds of millions to billions of images—across MacOS, Linux, and Windows.

Highlights

  • Extreme Scalability: Capable of processing 400 million images on a single CPU machine and scaling up to billions.
  • High Performance: Optimized C++ engine for speed even on low-resource hardware.
  • Hugging Face Integration: Ability to load and analyze datasets directly from the Hugging Face hub.
  • Privacy-First: Runs locally or on private cloud infrastructure, ensuring data remains in place.
  • Comprehensive Cleaning: Detects broken images, mislabels, and quality issues like blurriness and brightness.

Related

  • Project
  • Project
  • Project
  • Project
  • Project