OpenAI and Broadcom Unveil Jalapeño LLM Inference Chip

OpenAI and Broadcom Unveil Jalapeño LLM Inference Chip

OpenAI and Broadcom have unveiled Jalapeño, the first "Intelligence Processor" designed specifically for LLM inference. This custom accelerator aims to make advanced AI faster, more reliable, and more affordable by optimizing the hardware layer to match the specific requirements of frontier Large Language Models (LLMs).

High-Efficiency Architecture for LLM Inference

Jalapeño is a blank-slate design focused on LLM inference rather than being a general-purpose accelerator adapted from previous AI workloads. The architecture is optimized around the kernels, memory movement, networking, and serving patterns that are critical for frontier AI models.

Key technical goals and outcomes include:

  • Performance per Watt: Early testing indicates that the first-generation accelerator delivers performance per watt that is substantially better than current state-of-the-art hardware.
  • Hardware Utilization: The design reduces data movement and balances compute, memory, and networking resources to achieve realized utilization closer to the hardware's theoretical peak performance.
  • Full-Stack Integration: The chip is designed to work with all LLMs, informed by OpenAI's internal roadmap of models, kernels, serving systems, and product needs.

Engineering samples are currently running ML workloads in the lab at production target frequency and power, including the GPT-5.3-Codex-Spark model.

Accelerated Development Cycle

Jalapeño was developed from initial design to manufacturing tape-out in only nine months. OpenAI describes this as potentially the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors. This rapid timeline was made possible through:

  • Software-Hardware Co-development: Close collaboration between OpenAI engineering teams and Broadcom's silicon implementation expertise.
  • AI-Driven Design: OpenAI's own models were used to accelerate parts of the design and optimization process. -n- Industrialization Partners: Broadcom provided silicon implementation and networking technologies (including Tomahawk networking silicon), while Celestica provided board, rack, and system integration expertise.

Strategic Implications and Deployment

Jalapeño is the first iteration of a multi-generation compute platform. The strategy is to build a full-stack infrastructure—spanning chip architecture, kernels, memory systems, networking, and scheduling—to increase the abundance of compute.

Deployment Timeline

  • Initial Deployment: Scheduled for the end of 2026.
  • Scale: The platform is intended for deployment at "gigawatt scale" data centers in partnership with Microsoft and other partners.

Impact on End-User Experience

By increasing compute efficiency and lowering costs, OpenAI intends for these hardware improvements to translate directly into user-facing benefits: faster response times in ChatGPT, more complex task execution in Codex, and more affordable API products for developers and businesses.

Sources