dstackai/dstack

Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

What it solves

dstack 是一个统一的 GPU 供应与编排控制平面。它消除在不同 GPU 云、Kubernetes 集群和本地服务器之间管理计算资源的复杂性,提供一致的方式来处理开发、训练和推理。

How it works

用户部署 dstack 服务器和 CLI 来管理基础设施。系统通过配置“后端”连接各种 GPU 云或集群。用户在 YAML 配置文件中定义舰队、开发环境和任务的基础设施需求。运行 dstack apply 后,系统会自动处理资源供应、作业排队、自动扩缩、网络和卷管理,跨所有已连接的环境工作。

Who it’s for

需要将工作负载从本地开发扩展到分布式训练和模型部署的 AI 开发者和机器学习工程师,支持多种硬件加速器(NVIDIA、AMD、Google TPU 与 Tenstorrent)。

Highlights

  • Multi-cloud and Hybrid Support: Works across any GPU cloud, Kubernetes, and on‑prem clusters.
  • Detailed Resource Management: Supports fleets, dev environments, tasks, and services for different stages of the ML lifecycle.
  • uma own AI Agent Integration: Provides "skills" that allow AI agents (like Claude or Cursor) to manage fleets and submit workloads via the CLI.
  • Broad Hardware Compatibility: Out‑of‑the‑box support for NVIDIA, AMD, Google TPU, and Tenstorrent accelerators.