ml-energy/zeus

Measure and optimize the energy consumption of your AI applications!

What it solves

Zeus addresses the lack of visibility into the energy consumption of deep learning workloads. It provides tools to accurately measure and optimize the power usage of the hardware used to train and run AI models, helping developers reduce the environmental impact and operational costs of AI.

How it works

Zeus operates as a library and a daemon (zeusd) that abstracts the complexities of different hardware devices. It provides a programmatic interface and a CLI for energy monitoring, along with a collection of time and energy optimizers to improve efficiency. It supports a wide range of platforms, including CPU, DRAM, AMD GPUs, NVIDIA GPUs, Apple Silicon, and NVIDIA Jetson.

Who it’s for

ML engineers, researchers, and sustainability-focused AI developers who need to track the energy footprint of their models and optimize their training or inference pipelines for better energy efficiency.

Highlights

  • Broad Hardware Support: Supports NVIDIA GPUs, AMD GPUs, Apple Silicon, CPU, and DRAM.
  • PyTorch Ecosystem Project: Officially recognized as part of the PyTorch ecosystem.
  • Integrated Optimization: Includes tools like the Perseus optimizer for large model training.
  • Agent-Ready: Provides a portable "Agent Skill" to help AI coding agents automate energy measurement.
  • Research-Backed: Based on multiple peer-reviewed papers from NSDI, SOSP, and NeurIPS.

Related

  • Project
  • Project
  • Project
  • Project
  • Project