AI GlossaryㅍTechnical words in the news
Prompt Cache
A technique that saves the results of parts of a prompt an AI has already processed so it can reuse them instead of recomputing them, cutting both speed and cost
In plain words
A prompt cache lets an AI model reuse the results it already computed for the earlier part of a prompt, instead of processing that part all over again. It's like a kitchen that keeps a pot of broth already simmered, so instead of starting from scratch every time, the cook just pours in the ready-made broth and saves cooking time.
AI models read a prompt from start to end, in order, to generate a response. In long conversations, it's common for the earlier part of the text to stay the same while only the later part keeps changing. If the model had to re-read and recompute that unchanged earlier part every single time, it would waste both time and money. A prompt cache stores the computed results for that unchanged earlier part, and when the same earlier part shows up again, it skips recomputing it and picks up right where it left off.
This is also a big help when an agent rewinds to an earlier point in a task and restarts from there. Since everything up to that rewind point is identical, the cached results for that portion can be reused as is, and only the parts that differ afterward need to be computed fresh.
How it shows up in the news
The article says that for Shepherd, a tool that rewinds agent execution, "because the prompt prefix up to the branch point doesn't change, more than 95% of the prompt cache is reused on replay." Here, reusing the prompt cache means that even when execution is rewound to a specific point, the computation up to that point doesn't need to be redone — it doesn't mean the conversation content itself is being stored and remembered.
Try it yourself
Fix a long instruction (such as a role description or set of rules) as a system message, and send several different questions in a row within the same conversation. If you notice that the time it takes to get the first response gets shorter starting from the second question, that's a sign the earlier part is being cached and reused.
See also
Stories using this term
- Shepherd, open-source runtime for rewinding agent execution unveiledAI · 2026.08.09
- Claude Managed Agents update memory, domain controls, and session viewerAI · 2026.08.20
- Turing Award Winner Sutton: "Synthetic Data Can't Scale AI"AI · 2026.08.21
- Anthropic finds collaboration breaks down in agent swarm experimentsAI · 2026.08.17
- AWS Adds Open-Source Agent Skills for Bedrock Automated Reasoning PoliciesAI · 2026.08.09
- Cursor CEO's "must become multi-product company" remark resurfacesBusiness · 2026.08.15
