unitycatalog/unitycatalog

Open, Multi-modal Catalog for Data & AI

What it solves

Unity Catalog provides a universal, open-source catalog for managing data and AI assets. It solves the problem of fragmented governance by allowing users to govern and secure tabular data, unstructured assets, and AI models within a single interface, regardless of the underlying format or compute engine.

How it works

It operates as a multimodal interface with an open API (OpenAPI spec) and an OSS implementation. It supports multiple formats (such as Delta Lake, Apache Iceberg, Apache Hudi, Parquet, JSON, and CSV) and is compatible with Apache Hive's metastore API and Apache Iceberg's REST catalog API. This allows various compute engines to read and manage the data cataloged in Unity.

Who it’s for

This project is for data engineers, AI practitioners, and organizations that need a unified way to manage and secure their data and AI assets across different tools and engines.

Highlights

  • Multimodal Support: Manages tables, files, functions, and AI models.
  • Universal Compatibility: Works with multiple data formats and various leading compute engines.
  • Cros-Engine Governance: Provides unified governance for both structured and unstructured data and AI assets.
  • Open Standard: Based on an OpenAPI specification and hosted by the LF AI & Data Foundation.

Related

  • Project
  • Project
  • Project
  • Project
  • Project