Hugging Face and Google Cloud Strategic Partnership
Hugging Face and Google Cloud have announced a strategic partnership to enable companies to build and customize their own AI using open models on Google Cloud infrastructure. This collaboration aims to simplify the process of deploying open models across various Google Cloud services and improve the performance and security of the Hugging Face Hub.
Optimized Model Delivery via CDN Gateway
Google Cloud customers will benefit from a new CDN Gateway for Hugging Face repositories. This gateway is designed to reduce download times and increase supply chain robustness by caching Hugging Face models and datasets directly on Google Cloud.
Built using Hugging Face Xet optimized storage and data transfer technologies combined with Google Cloud's storage and networking capabilities, the CDN Gateway provides faster time-to-first-token and simplified model governance for users across Vertex AI, GKE, Cloud Run, and Compute Engine VMs.
Integration Across Google Cloud AI Services
Open models from Hugging Face are integrated into several Google Cloud AI services to provide flexible deployment options:
- Vertex AI: Popular open models are available for deployment via Model Garden.
- GKE AI/ML: A model library is available for customers requiring greater control over their AI infrastructure, including pre-configured environments maintained by Hugging Face.
- Cloud Run GPUs: These enable serverless deployments of open models for AI inference workloads.
Enhancements for Hugging Face Customers
The partnership extends capabilities to Hugging Face users through the following initiatives:
Infrastructure and Hardware Acceleration
Hugging Face Inference Endpoints will integrate more Google Cloud instances, offering new instance types and price reductions. Additionally, Hugging Face is working to make Google's seventh-generation TPUs (Tensor Processing Units) as easy to use as GPUs for open models through native support in Hugging Face libraries.
Security and Model Governance
Hugging Face will leverage Google's security technology to secure the millions of models, datasets, and Spaces on the Hugging Face Hub. This security effort is powered by VirusTotal, Google Threat Intelligence, and Mandiant.
Deployment Workflow
The goal is to streamline the transition from the Hugging Face Hub to Google Cloud. This includes simplifying the deployment of both public and private models hosted in Enterprise organizations from Hugging Face model pages directly to Vertex Model Garden or GKE.
Strategic Goals
The partnership focuses on a future where companies maintain full control over their AI by hosting open models within their own secure infrastructure. This vision is supported by the developments in Vertex AI Model Garden, Google Kubernetes Engine, and Hugging Face Inference Endpoints.