data-infra/cube-studio

cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持

What it solves

CubeStudio is a cloud-native, one-stop AI platform designed to eliminate the repetitive engineering overhead that algorithm engineers face. It streamlines the entire AI lifecycle—from data preparation and labeling to development, training, evaluation, and deployment—reducing the time spent on infrastructure management (like requesting GPUs, configuring environments, and managing data flows) so developers can focus on model building.

How it works

Built on a Kubernetes foundation, CubeStudio manages heterogeneous computing resources (including NVIDIA, AMD, and various domestic Chinese AI chips like Ascend and Cambricon). It provides a layered architecture:

  • Infrastructure: Handles GPU/NPU scheduling, vGPU virtualization, and RDMA high-speed networking.
  • Development: Offers browser-based IDEs (JupyterLab, VSCode) and image management for consistent environments.
  • Data Management: Integrates data exploration (SQLLab), dataset management, and a multi-modal labeling platform with AI-assisted automation.
  • Training: Uses a drag-and-drop Pipeline orchestrator to build DAGs for ML/DL workflows, supporting distributed training (PyTorch, TensorFlow, DeepSpeed) and hyperparameter search.
  • Deployment: Manages model versioning and provides a zero-code deployment path to inference services with support for A/B testing, canary releases, and auto-scaling.

Who it’s for

  • AI Engineers and Data Scientists: Who need a unified environment for the full ML lifecycle without managing low-level infrastructure.
  • Enterprises: Seeking a private, sovereign AI platform that supports domestic hardware and complies with data security requirements.
  • MLOps Teams: Who need to manage multi-tenant resource allocation, billing, and monitoring across multiple clusters.

Highlights

  • Heterogeneous Hardware Support: Native compatibility with a wide range of domestic AI chips (Ascend, Cambricon, Hygon, etc.) and ARM/x86 architectures.
  • Full-Stack MLOps: Covers everything from data labeling and ETL to SFT (Supervised Fine-Tuning) and RLHF for LLMs.
  • Drag-and-Drop Pipelines: Simplifies complex workflow orchestration via a visual interface.
  • LLM-Ready: Includes built-in support for vLLM, Ollama, and MindIE for high-performance large model inference and fine-tuning.
  • Enterprise-Grade Governance: Features multi-tenancy, RBAC, resource quotas, and detailed metering/billing.

Related

  • Project
  • Project
  • Project
  • Project
  • Project