OpenAI updates prompt caching for GPT-6 API
Shared prompt prefixes reused within a 30-minute window now qualify for input token discounts of up to 90%.
- Developers can adjust reasoning effort and tool availability mid-conversation via configuration_update without invalidating the cache.
- A diagnostics API and tool detail reasons for cache misses and the number of affected tokens.
- Explicit cache breakpoints allow applications to choose specific prompt prefixes to preserve.
- Cache prewarming prepares shared instructions and tool definitions before user requests arrive.
- A Prompt Caching Dashboard tracks hit rates and token composition over time.
Developers building multi-turn agents on GPT-6 can lower API costs and latency while dynamically modifying session configurations.

Sources
Read this as text
Back to the AI news