METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryKTechnical words in the news

KV Cache

A temporary memory space that stores what an AI has already read so it doesn't have to recompute it

In plain words

A KV cache is a memory space where an AI temporarily stores calculations it has already done, so that when it keeps writing text, it doesn't have to reread everything from the beginning.

Imagine someone reading a long book aloud. If they had to reread every previous page from scratch before reading each new sentence, they'd get slower and slower as the book went on. Instead, it's much faster to keep a running set of notes on what's been read so far and glance at them when needed. When an AI generates a response one piece at a time, it does something similar: it stores the calculation results for the words that came before, like a stack of notes, and reuses them. This stack of notes is the KV cache.

The problem is that as a conversation or document gets longer, this stack of notes grows thicker too. If the desk where the notes pile up — that is, the memory space on the graphics card — runs out of room, even a powerful AI struggles to handle long text. That's why attempts to compress models to free up more of this memory space come up so often.

How it shows up in the news

One article explains the reasoning behind compressing a model, saying it "leaves enough room to run the KV cache, a perception encoder, and a speculative decoding drafter at the same time, even on a 24GB or 32GB-class GPU." Another article points out that "in practice, the KV cache adds on top of that. With long context, the cache can end up as large as the weights themselves." A common misconception is thinking that securing enough space for just the model file is enough to run it — but in reality, as a conversation gets longer, this cache eats up additional memory on its own, which is easy to overlook.

Try it yourself

Ask a chatbot a short question and notice how fast it responds. Then paste in a very long document and ask a question about it, comparing how long it takes for the response to start. You'll notice that the longer the conversation (or pasted document), the more notes the AI has to reference, and the longer it takes before the first response appears.

See also

Stories using this term

Browse every entry