awslabs/data-on-eks

DoEKS is a tool to build, deploy and scale Data Platforms on Amazon EKS

What it solves

Data on EKS (DoEKS) simplifies the deployment and scaling of data platforms on Amazon EKS. It addresses the complexity of managing the Kubernetes ecosystem and selecting the right configurations for big data workloads, allowing users to quickly build production-ready clusters and conduct proof-of-concepts.

How it works

DoEKS provides a set of opinionated open-source blueprints that integrate various data tools, Kubernetes operators, and AWS Data Analytics managed services. These blueprints automate the deployment of frameworks for distributed data processing, real-time stream processing, and workflow orchestration.

Who it’s for

Data engineers and platform architects who need to run scalable data workloads—such as Spark, Flink, Kafka, and Airflow—on Kubernetes using Amazon EKS.

Highlights

  • Data Processing: Blueprints for Apache Spark, Amazon EMR on EKS, and Ray.
  • Streaming Platforms: Support for Apache Flink and Apache Kafka (via Strimzi).
  • Orchestration: Deployment patterns for Apache Airflow and Argo Workflows.
  • Databases & Query Engines: Integration with Trino, Pinot, and ClickHouse.

Related

  • Project
  • Project
  • Project
  • Project
  • Project