AI GlossaryㅋTechnical words in the news
Cache Hit
When an AI receives the exact same input it processed before and reuses the stored result instead of recalculating it from scratch
In plain words
A cache hit happens when part of what you send an AI has already been processed before, so the system pulls the saved result instead of computing everything again from the start. The opposite case—brand-new content the system hasn't seen—is called a cache miss, which requires a full recalculation.
A restaurant is a good analogy. If you prep the same ingredients ahead of time, the next customer's order comes out much faster. This is exactly what happens with instructions that get attached to every request (system prompts), or when you ask several questions in a row about the same document. Since the earlier part matches what came before, it doesn't need to be recalculated—it's simply reused.
Because of this reuse, AI companies charge much less for the portion that hits the cache compared to a cache miss. As a result, the more you repeatedly reference the same conversation or document, the lower your overall cost becomes.
How it shows up in the news
Articles describe things like: "the price for cache-hit input is $0.003625 per million tokens, about one-hundredth the price of a cache miss." One point that's easy to confuse: a cache hit has nothing to do with the model getting smarter or the answer quality changing. It's purely a pricing policy and technical mechanism that reduces the computational cost of reprocessing the same input.
Try it yourself
- On a service that uses an API, send the same request—including an identical system prompt or long document—twice in a row.
- Check whether the usage information returned with the response separately shows the number of cache-hit tokens.
- Compare the cost or response speed between the first and second requests to feel the effect of the cache hit firsthand.
See also
Stories using this term
- DeepSeek-V4-Pro moves to general availability, API pricing revealedAI · 2026.08.13
- Microsoft's new coding model falls short of DeepSeek on both price and performanceAI · 2026.08.12
- Anthropic Captures 65% of Vercel AI Gateway Revenue with Just 30% of TokensBusiness · 2026.08.19
- OpenAI cuts GPT-5.6 Sol API pricing by over 20% for three monthsAI · 2026.08.22
- Stripe Finalizes Acquisition of AI Gateway OpenRouter for Over $7 BillionBusiness · 2026.08.17
- DeepSeek Forms Harness Team, Takes Aim at Claude CodeBusiness · 2026.08.13
