AWS and OpenAI Strategic Partnership

AWS and OpenAI have announced a multi-year strategic partnership to provide OpenAI with immediate access to AWS infrastructure to scale its core AI workloads. This $38 billion agreement, spanning seven years, allows OpenAI to expand its compute capacity using AWS's large-scale infrastructure to support the training of next-generation models and the serving of inference for ChatGPT.

Compute Infrastructure and Scaling

OpenAI will utilize AWS compute consisting of hundreds of thousands of state-of-the-art NVIDIA GPUs, with the capacity to expand to tens of millions of CPUs. This infrastructure is specifically designed to support the rapid scaling of agentic workloads and the general demand for computing power required by frontier AI models.

Hardware and Architecture

The infrastructure deployment features an architectural design optimized for AI processing efficiency. Key technical specifications include:

  • Hardware: The deployment utilizes Amazon EC2 UltraServers to cluster NVIDIA GB200s and GB300s.
  • Connectivity: By clustering these GPUs on the same network, the system achieves low-latency performance across interconnected systems.
  • Deployment Timeline: All initial capacity is targeted for deployment by the end of 2026, with options for further expansion into 2027 and beyond.

Strategic Objectives and Model Deployment

This partnership aims to strengthen the compute ecosystem necessary for power the next era of AI. According to OpenAI CEO Sam Altman, "Scaling frontier AI requires massive, reliable compute."

Integration with Amazon Bedrock

Prior to this strategic partnership, OpenAI open weight foundation models were made available on Amazon Bedrock. This has already led to the thousands of customers—including Peloton, Thomson Reuters, Comscore, Bystreet, Triomics, and Verana Health—utilizing OpenAI models for tasks such as:

  • Agentic workflows
  • Coding
  • Coding and scientific analysis
  • Mathematical problem-solving

Operational Scale and Security

AWS provides the scale and security required for OpenAI's vast workloads. AWS has experience managing clusters topping 500,000 chips, which ensures the reliability and reliability of the infrastructure used to run ChatGPT and train future models.

Sources