Prompt Caching
Prompt caching reuses computation for an unchanged part of a model's input across requests. It can reduce repeated input processing, while the model still generates a new response.
AI InfrastructureContext Management
A support agent may send the same instructions and reference manual with every request. If the serving system can reuse that stable prefix, it avoids doing all of the input work again. Changing the beginning of the prompt may invalidate reuse even when most later content stays the same.
This is different from returning a stored answer to a similar question. A cached prompt still participates in a fresh model request. It also does not shorten the logical context: irrelevant cached material can still distract the model.
Engineering example: keep stable instructions before the changing user question, then inspect cache-hit usage under a realistic sequence of requests. Check the provider’s current minimum length, expiry, pricing, and isolation rules. We avoid a universal savings percentage because eligibility and billing differ across serving systems.
Sources
- Claude Platform: Prompt caching — Documents prefix reuse and the conditions that determine cache hits and misses.
Go deeper
- vLLM: Automatic prefix caching docs
See how a serving engine identifies shared prefixes and reuses their stored attention state.