apache/gravitino
World's most powerful open data catalog for building a high-performance, geo-distributed and federated metadata lake.
What it solves
Apache Gravitino provides a unified way to manage metadata for both data and AI assets across different sources, types, and geographical regions. It eliminates the need to manage multiple disparate metadata catalogs up close, allowing users to access and govern data and AI assets in a federated manner.
How it works
Gravitino acts as a federated metadata lake. It uses connectors to integrate directly with underlying systems (like Hive, MySQL, MariaDB, HDFS, and S3) so that changes in the source systems are immediately reflected. It provides a single API and model to manage these diverse sources, and integrates with query engines like Trino and Spark without requiring changes to SQL dialects.
Who it’s for
Data engineers, AI practitioners, and architects who need to manage metadata across multi-cloud or hybrid setups, synchronize metadata across regions, and implement unified governance (access control and auditing) for data and AI assets.
Highlights
- Unified Metadata Management: A single API for diverse sources including Hive, MySQL, MariaDB, HDFS, and S3.
- End-to-End Governance: Integrated access control, auditing, and discovery across all assets.
- Geo-Distribution: Support for sharing metadata across different regions and clouds.
- AI Asset Management: Support for tracking AI models and features (currently a work in progress).
- Native REST Catalogs: Provides native Iceberg and Lance REST catalog services.
- Multi-Engine Compatibility: Seamless integration with engines like Trino and Spark.
Related
- Project
- Project
- Project
- Project
- Project