awslabs/data-on-eks
DoEKS is a tool to build, deploy and scale Data Platforms on Amazon EKS
What it solves
Data on EKS (DoEKS) simplifies the deployment and scaling of data platforms on Amazon EKS. It addresses the complexity of managing the Kubernetes ecosystem and selecting the right configurations for big data workloads, allowing users to quickly build production-ready clusters and conduct proof-of-concepts.
How it works
DoEKS provides a set of opinionated open-source blueprints that integrate various data tools, Kubernetes operators, and AWS Data Analytics managed services. These blueprints automate the deployment of frameworks for distributed data processing, real-time stream processing, and workflow orchestration.
Who it’s for
Data engineers and platform architects who need to run scalable data workloads—such as Spark, Flink, Kafka, and Airflow—on Kubernetes using Amazon EKS.
Highlights
- Data Processing: Blueprints for Apache Spark, Amazon EMR on EKS, and Ray.
- Streaming Platforms: Support for Apache Flink and Apache Kafka (via Strimzi).
- Orchestration: Deployment patterns for Apache Airflow and Argo Workflows.
- Databases & Query Engines: Integration with Trino, Pinot, and ClickHouse.
Related
- Project
- Project
- Project
- Project
- Project