The overlooked result
A persistent agent may make dozens of requests while carrying forward the same instructions, tool definitions and project context. Reprocessing all of that text on every step wastes time and money. OpenAI’s revised prompt caching for GPT-6 reuses eligible shared prefixes and discounts cached input tokens by up to 90%.
The new system supports reuse of shared prefixes within a 30-minute window. That does not mean every prompt automatically gets a 90% discount: the reused part must be eligible, stable and actually hit the cache. For a high-volume application, knowing which portion was cached matters more than the headline percentage.
What builders can change
OpenAI added a Prompt Caching Dashboard that separates cached from uncached input over time. A diagnostics tool compares a request with a recent response and identifies changes to settings, tools or input that may have broken reuse. Explicit breakpoints give developers more control over where the stable prefix ends.
GPT-6 also allows an application to change reasoning effort between responses without discarding earlier cached context, provided it follows the supported configuration path. Tool schemas and ordering should remain stable; an agent can limit callable tools without needlessly rewriting the shared prefix. Prewarming lets known context be processed before a user waits on the first response.
The measured payoff
OpenAI cites GitHub Copilot’s report that, across billions of requests, the share of prompt tokens needing fresh processing fell by more than 50% against its earlier baseline. Other customer examples on the launch page report higher cache-hit rates and lower costs after tuning breakpoints and diagnosing misses.
The lesson is specific: benchmark the full workflow, record cache-hit rate and task quality, then change the prompt layout. A cache hit lowers input-processing cost; it does not repair a weak agent or eliminate the need to check its work. For repetitive, long-context applications, though, this can be a larger practical gain than a small model-score improvement.
Source published 2026-09-22. Coverage is based on the maker’s announcement and demonstration.
