data-infra/cube-studio
cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持
What it solves
CubeStudio addresses the complexities of managing machine learning lifecycles by providing a unified, cloud-native platform. It solves the challenges of resource allocation, user permission management, and multi-hardware scheduling in large-scale AI environments.
How it works
The platform utilizes a cloud-native architecture to manage various computing resources and services. It employs a role-based access control (RBAC) system for user and project management and supports diverse hardware through advanced scheduling. It integrates with various storage solutions and can operate across multiple Kubernetes clusters, including edge and serverless environments.
Who it’s for
It is designed for organizations and development teams that require a centralized platform to manage machine learning tasks, including development, training, and inference services, across heterogeneous hardware and distributed clusters.
Highlights
- Diverse Hardware Support: Supports CPU, GPU (T4, V100, A100), ARM64, and various domestic GPUs (Hygon DCU, Huawei NPU, Cambricon MLU, etc.).
- Comprehensive Resource Management: Includes project/user management, RBAC, and detailed metering/billing for development, training, and inference resources.
- Flexible Deployment: Supports multiple Kubernetes clusters, edge computing nodes, and serverless modes (Tencent Cloud and Alibaba Cloud).
- Extensive Storage Integration: Supports NFS, CFS, OSS, NAS, COS, GlusterFS, CephFS, and S3/MinIO.
- Advanced Networking: Supports HTTPS, reverse proxies, and various network tunneling methods.