AI GlossaryㅋTechnical words in the news
Cached Prompt
A way of processing prompts that reuses previously computed results instead of recalculating them, so it costs less.
In plain words
A cached prompt means that when you send the same content to an AI again, instead of recalculating everything from scratch, it pulls up results that were already processed before. Think of a cafe: instead of grinding beans and boiling water every time a customer arrives, the barista just reheats coffee that's already brewed. The taste (result) is nearly the same, but the effort and cost of making it drop a lot.
AI agents often keep reloading the same instructions or background information over and over until a task is finished. If the overlapping parts were recalculated every single time, the time and cost would balloon. With cached prompts, the overlapping parts reuse stored results, and only the new parts get calculated. So even for the same amount of work, the price per unit ends up much lower than a first-time calculation.
As a result, cost often doesn't rise nearly as much as usage does, even when usage multiplies several times over. The more repetitive an AI agent's work is, the bigger the share of cached prompts becomes, and the more slowly actual billed costs climb compared to the jump in usage.
How it shows up in the news
The article notes that on OpenRouter, roughly 70% of all agent tokens come from cached prompts. That's why a 14x increase in an agent's token consumption doesn't mean costs also jump 14x — an easy point to misread, since caching makes usage growth and cost growth move at different speeds.
Try it yourself
You can feel the difference by sending the same instructions or long background text repeatedly through an AI chatbot or API. For example, if you lay out the same long document as background and just change the questions you ask about it, you'll notice faster responses or lower costs compared to swapping in completely new content each time. The more overlap there is, the better caching works.
See also
Stories using this term
- Grok 4.6 unveiled: top-tier performance at half the priceAI · 2026.08.13
- NVIDIA releases Switchyard, an LLM routing proxy for coding agentsAI · 2026.08.20
- Coding-focused stealth model 'Ox Alpha' appears free on OpenRouterAI · 2026.08.22
- OpenAI Blocks Cluster of ChatGPT Accounts Behind Russian Influence CampaignAI · 2026.08.26
- OpenAI Expands DevDay to 8 Cities Including SeoulBusiness · 2026.08.19
- OpenAI Investor Thrive Holdings Raises $2 Billion at $12 Billion ValuationBusiness · 2026.08.13
