OpenAI Prompt Caching API Release
OpenAI has launched Prompt Caching, a feature that allows developers to reduce costs and latency by reusing recently seen input tokens. This is particularly beneficial for applications involving long, multi-turn conversations or repeated edits to a large codebase.
Prompt Caching Pricing and Model Availability
Prompt Caching is automatically applied to the latest versions of GPT-4o, GPT-4o mini, o1-preview, and o1-mini, including their fine-tuned versions. Cached input tokens are priced at a 50% discount compared to uncached input tokens.
| Model | Uncached Input Tokens | Cached Input Tokens | Output Tokens |
|---|---|---|---|
| gpt-4o-2024-08-06 | $2.50 | $1.25 | $10.00 |
| GPT-4o fine-tuning | $3.75 | $1.875 | $15.00 |
| gpt-4o-mini-2024-07-18 | $0.15 | $0.075 | $0.60 |
| GPT-4o mini fine-tuning | $0.30 | $0.15 | $1.20 |
| o1-preview | $15.00 | $7.50 | $60.00 |
| o1-mini | $3.00 | $1.50 | $12.00 |
Technical Implementation and Cache Logic
Prompt Caching is automatically applied to prompts longer than 1,024 tokens. The system caches the longest prefix of a prompt that has been previously computed, starting at a minimum of 1,024 tokens and increasing in 128-token increments.
Developers do not need to make changes to their API integration to benefit from the discount, as the system automatically identifies and applies the caching for common prefixes.
Cache Monitoring and Lifecycle
API responses now include a cached_tokens value within the usage field, allowing developers to monitor cache usage.
Caches are managed based on a period of inactivity:
- Caches are typically cleared after 5 to 10 minutes of inactivity.
- Caches are always removed within one hour of the last use.
Privacy and Security
Prompt caches are not shared between organizations, ensuring that cached data remains isolated. This feature is subject to OpenAI's Enterprise privacy commitments.
Sources
- OriginalPrompt Caching in the API