dstackai/dstack

Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

What it solves

dstack は GPU のプロビジョニングとオーケストレーションを統合したコントロールプレーンです。異なる GPU クラウド、Kubernetes クラスター、オンプレミスサーバー間でコンピュートリソースを管理する複雑さを取り除き、開発・トレーニング・推論を一貫した方法で扱えるようにします。

How it works

ユーザーは dstack サーバーと CLI をセットアップしてインフラを管理します。システムは「バックエンド」を設定して各種 GPU クラウドやクラスターに接続します。ユーザーは YAML 設定ファイルでフリート、開発環境、タスクのインフラ要件を定義します。dstack apply を実行すると、システムは自動的にプロビジョニング、ジョブキューイング、オートスケーリング、ネットワーキング、ボリューム管理を行い、接続されたすべての環境で動作します。

Who it’s for

ローカル開発から分散トレーニング、モデルデプロイまでワークロードをスケールさせる必要がある AI 開発者や ML エンジニア向けです。対応ハードウェアは NVIDIA、AMD、Google TPU、Tenstorrent です。

Highlights

  • Multi-cloud and Hybrid Support: Works across any GPU cloud, Kubernetes, and on‑prem clusters.
  • Detailed Resource Management: Supports fleets, dev environments, tasks, and services for different stages of the ML lifecycle.
  • uma own AI Agent Integration: Provides "skills" that allow AI agents (like Claude or Cursor) to manage fleets and submit workloads via the CLI.
  • Broad Hardware Compatibility: Out‑of‑the‑box support for NVIDIA, AMD, Google TPU, and Tenstorrent accelerators.