Hugging Face and NVIDIA Launch Training Cluster as a Service
Hugging Face and NVIDIA have launched Training Cluster as a Service, a collaborative initiative designed to make large-scale GPU clusters accessible to research organizations, universities, and companies globally. This service aims to bridge the compute gap by allowing organizations to request and pay for high-performance GPU capacity specifically for the duration of their training runs.
On-Demand GPU Cluster Accessibility
Training Cluster as a Service provides a flexible procurement model for AI compute, enabling any of the 250,000 organizations on Hugging Face to request specific GPU cluster sizes based on their immediate needs. The primary goal is to enable the development of foundational models across diverse domains by removing the financial and logistical barriers associated with securing massive compute resources.
Technical Architecture and Workflow
The service integrates infrastructure from NVIDIA with the developer ecosystem of Hugging Face through the following components:
- NVIDIA Cloud Partners: Provide the physical capacity for the latest accelerated computing hardware, including NVIDIA Hopper and NVIDIA GB200, located in regional datacenters.
- NVIDIA DGX Cloud: Acts as the centralized management layer for these regional resources.
- NVIDIA DGX Cloud Lepton: A tool announced at GTC Paris that provides researchers with access to provisioned infrastructure and manages the scheduling and monitoring of training runs.
- Hugging Face Libraries: Open-source libraries and developer resources are used to initiate and manage the training workloads.
Once a request is submitted via hf.co/training-cluster, Hugging Face and NVIDIA collaborate to source, price, and provision the cluster based on the organization's requirements for size, region, and duration.
Real-World Research Applications
Several organizations are already utilizing Training Cluster as a Service to advance specialized AI research:
- TIGEM (Telethon Institute of Genomics and Medicine): Using the service to predict the effect of pathogenic variants and explore drug repositioning for rare genetic diseases.
- Numina: A non-profit building open-source AI for mathematical reasoning to create open alternatives to closed-source models like DeepMind's AlphaProof.
- Mirror Physics: A startup developing high-fidelity chemical models at scale for chemistry and materials science.
Broader NVIDIA and Hugging Face Integration
Beyond the training cluster service, the collaboration includes several other technical integrations announced at GTC Paris:
- NVIDIA NIM: NVIDIA AI customers can now deploy over 100,000 Hugging Face models using NVIDIA Inference Microservices (NIM).
- NVIDIA Cosmos Predict-2: Enables Hugging Face users to build custom Physical AI models.
- NVIDIA Isaac GR00T N1.5: Now available on Hugging Face to power the development of humanoid robots.