open-edge-platform/datumaro

Dataset Management Framework, a Python library and a CLI tool to build, analyze and manage Computer Vision datasets.

What it solves

Datumaro は、AIモデルのトレーニング用にデータセットを準備する複雑なプロセスを簡素化します。データセットの形式が断片化し、アノテーションが不一致であり、モデルをトレーニングする前に厳格なデータクリーニングと分割が必要であるという問題に対処処します。

How it works

フレームワークと CLI ツールを通じて、データセット管理の中央ハブとして機能します。業界標準の形式(COCO, Pascal VOC, YOLO など)を幅広く読み書きできるため、ユーザーはこれらの形式間でデータセットを変換できます。変換だけでなく、n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/A, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a,n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/a, n/ industry-standard formats (such as COCO, Pascal VOC, and YOLO), allowing users to convert datasets between these formats. Beyond conversion, it provides tools to merge datasets, filter out unwanted data based on specific criteria, transform annotations (e.g., converting polygons to masks), and split data into training, validation, and test sets while maintaining label distributions.

Who it’s for

コンピュータビジョン用のデータセットを構築、変換、および分析する必要があるデータサイエンティストや ML エンジニア。

Highlights

  • Multi-format support: Seamlessly converts between numerous formats including COCO, YOLO, Cityscapes, and ImageNet.
  • Dataset building: Tools for merging datasets and filtering images or annotations based on custom criteria.
  • Intelligent splitting: Supports random splits or task-specific splits that preserve label and attribute distributions.
  • Quality control: Includes features for error checking, annotation validation, and comparison with model inference results.
  • Dataset statistics: Generates image mean, standard deviation, and annotation statistics.