METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅋTechnical words in the news

Cache-to-Cache

A method where two AI models pass information by directly linking their internal computed values, instead of exchanging written sentences

In plain words

Cache-to-Cache is a way of connecting two AI models where, instead of one writing text for the other to read, they hand over internal computed values directly. The common approach until now was for a larger model to output an answer as text, which a smaller model would then read from scratch and try to understand — like writing a letter and passing it to a coworker at the next desk. The problem is that writing the letter takes time, and the act of putting things into words flattens out subtle nuances along the way.

Cache-to-Cache removes this letter-writing step. It lets a model pass along the intermediate computed values it holds in its "head" before turning them into sentences, sending them directly to another model. The receiving model gets the original, rich judgment itself rather than a single flattened sentence, so it doesn't need to guess at what was meant. Research found this approach was both more accurate and much faster than connecting models through text.

Around the same time, a company called Mostic went a step further, trying to connect two models not by passing intermediate values back and forth, but by finding a conversion table between the fixed numbers (weights) each model holds after training. The approaches differ, but the goal is the same: letting models share each other's judgments without explaining themselves in words.

How it shows up in the news

In the article, this term — which is also the title of a research paper — appears alongside experimental results showing it was "3.1–5.4% more accurate and on average 2.5 times faster than connecting models through text." A common misunderstanding is that this isn't a technique for making a single model smarter — it's a connection method for linking multiple existing models more efficiently.

See also

Stories using this term

Browse every entry