dstackai/dstack
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
What it solves
dstack 是一個統一的 GPU 配置與編排控制平面。它消除在不同 GPU 雲端、Kubernetes 叢集與本地伺服器之間管理運算資源的複雜性,提供一致的方式來處理開發、訓練與推論。
How it works
使用者部署 dstack 伺服器與 CLI 以管理基礎設施。系統透過設定「後端」連接各種 GPU 雲端或叢集。使用者在 YAML 設定檔中定義艦隊、開發環境與任務的基礎設施需求。執行 dstack apply 後,系統會自動處理資源配置、工作排隊、自動擴縮、網路與磁碟管理,跨所有已連線的環境運作。
Who it’s for
需要將工作負載從本機開發擴展至分散式訓練與模型部署的 AI 開發者與機器學習工程師,支援多種硬體加速器(NVIDIA、AMD、Google TPU 與 Tenstorrent)。
Highlights
- Multi-cloud and Hybrid Support: Works across any GPU cloud, Kubernetes, and on‑prem clusters.
- Detailed Resource Management: Supports fleets, dev environments, tasks, and services for different stages of the ML lifecycle.
- uma own AI Agent Integration: Provides "skills" that allow AI agents (like Claude or Cursor) to manage fleets and submit workloads via the CLI.
- Broad Hardware Compatibility: Out‑of‑the‑box support for NVIDIA, AMD, Google TPU, and Tenstorrent accelerators.