METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅍTechnical words in the news

Prompt Cache

A technique that saves the results of parts of a prompt an AI has already processed so it can reuse them instead of recomputing them, cutting both speed and cost

In plain words

A prompt cache lets an AI model reuse the results it already computed for the earlier part of a prompt, instead of processing that part all over again. It's like a kitchen that keeps a pot of broth already simmered, so instead of starting from scratch every time, the cook just pours in the ready-made broth and saves cooking time.

AI models read a prompt from start to end, in order, to generate a response. In long conversations, it's common for the earlier part of the text to stay the same while only the later part keeps changing. If the model had to re-read and recompute that unchanged earlier part every single time, it would waste both time and money. A prompt cache stores the computed results for that unchanged earlier part, and when the same earlier part shows up again, it skips recomputing it and picks up right where it left off.

This is also a big help when an agent rewinds to an earlier point in a task and restarts from there. Since everything up to that rewind point is identical, the cached results for that portion can be reused as is, and only the parts that differ afterward need to be computed fresh.

How it shows up in the news

The article says that for Shepherd, a tool that rewinds agent execution, "because the prompt prefix up to the branch point doesn't change, more than 95% of the prompt cache is reused on replay." Here, reusing the prompt cache means that even when execution is rewound to a specific point, the computation up to that point doesn't need to be redone — it doesn't mean the conversation content itself is being stored and remembered.

Try it yourself

Fix a long instruction (such as a role description or set of rules) as a system message, and send several different questions in a row within the same conversation. If you notice that the time it takes to get the first response gets shorter starting from the second question, that's a sign the earlier part is being cached and reused.

See also

Stories using this term

Browse every entry