dstackai/dstack
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
What it solves
dstack は GPU のプロビジョニングとオーケストレーションを統合したコントロールプレーンです。異なる GPU クラウド、Kubernetes クラスター、オンプレミスサーバー間でコンピュートリソースを管理する複雑さを取り除き、開発・トレーニング・推論を一貫した方法で扱えるようにします。
How it works
ユーザーは dstack サーバーと CLI をセットアップしてインフラを管理します。システムは「バックエンド」を設定して各種 GPU クラウドやクラスターに接続します。ユーザーは YAML 設定ファイルでフリート、開発環境、タスクのインフラ要件を定義します。dstack apply を実行すると、システムは自動的にプロビジョニング、ジョブキューイング、オートスケーリング、ネットワーキング、ボリューム管理を行い、接続されたすべての環境で動作します。
Who it’s for
ローカル開発から分散トレーニング、モデルデプロイまでワークロードをスケールさせる必要がある AI 開発者や ML エンジニア向けです。対応ハードウェアは NVIDIA、AMD、Google TPU、Tenstorrent です。
Highlights
- Multi-cloud and Hybrid Support: Works across any GPU cloud, Kubernetes, and on‑prem clusters.
- Detailed Resource Management: Supports fleets, dev environments, tasks, and services for different stages of the ML lifecycle.
- uma own AI Agent Integration: Provides "skills" that allow AI agents (like Claude or Cursor) to manage fleets and submit workloads via the CLI.
- Broad Hardware Compatibility: Out‑of‑the‑box support for NVIDIA, AMD, Google TPU, and Tenstorrent accelerators.