OpenAI GPT-6 Prompt Caching Updates

OpenAI has launched an improved prompt caching system for the GPT-6 family of models to support persistent agents performing complex, multi-turn tasks. This system reduces response times and provides developers with discounts of up to 90% on cached input tokens by reusing shared context across API requests.

Improved Cache Hit Rates and Pricing

GPT-6 implements a caching system that delivers higher cache hit rates by default. The system provides cache discounts for eligible shared prefixes that are reused within a 30-minute window.

Monitoring and Diagnostics Tools

OpenAI has introduced two new tools to help developers manage and optimize their cache performance:

  • Prompt Caching Dashboard: A tool that allows developers to track hit rates over time and use an input composition chart to compare cached versus uncached tokens.
  • Prompt Caching Diagnostics Tool: A utility for investigating unexpected cache misses. It allows developers to compare a request with a recent response to identify changes in the model, tools, settings, or input that prevented reuse. The tool provides the estimated number of affected tokens to help assess the impact of a cache miss.

Optimization Strategies for GPT-6 Caching

Developers can use several new controls to maximize cache hit rates and reduce latency:

Explicit Cache Breakpoints

Developers can now choose which prompt prefixes to reuse using explicit cache breakpoints, allowing for more granular control over what is cached.

Dynamic Reasoning Effort

On GPT-6 models, developers can change the reasoning effort between responses without breaking the cache. This is achieved by appending a configuration_update while leaving the request-level reasoning effort unchanged, allowing for adjustments in reasoning intensity based on the task difficulty without losing reusable context.

Stable Tool and Instruction Management

To preserve the cache as tools and instructions evolve, OpenAI recommends keeping tool definitions, schemas, and ordering stable. Developers are advised to:

  • Use allowed_tools to limit callable tools rather than removing definitions.
  • Set tool_choice to none when no tools are needed.
  • Use developer messages to append new instructions at the end of the context to override older ones.

Cache Prewarming

Prewarming allows applications to prepare known context—such as shared instructions, tool definitions, or reference material—ahead of time. This moves the processing of that context out of the user's wait time, reducing initial response latency.

Sources