METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅎTechnical words in the news

Sparse Attention

A method that cuts computation by having an AI model focus only on the important parts of its input instead of comparing everything against everything else.

In plain words

Sparse attention is a way for an AI model to read a long text or conversation without re-scanning everything that came before word by word. It's like reviewing a thick meeting transcript not by rereading it from start to finish, but by checking the table of contents or an index and flipping only to the relevant pages.

In the standard approach, every word in a sentence has to be compared against every other word, so the computation snowballs as the text gets longer. Sparse attention keeps only the truly important relationships and skips the rest. This lets a model handle much longer text while saving a lot of compute and response time.

However, the rule for deciding which parts to keep and which to skip differs from model to model. How well that rule is designed determines the balance between speed and accuracy.

How it shows up in the news

The article cites MiniMax's own sparse attention technique, 'MSA (MiniMax Sparse Attention),' as the reason its M3 model can respond quickly even while handling a long context of up to one million tokens. By computing only the necessary key-value blocks, this technique reportedly allowed Modular Cloud to handle traffic on the order of billions of tokens per minute. Contrary to a common misconception, sparse attention isn't a technique that trades away model performance (accuracy) just to gain speed—its design actually focuses on cutting unnecessary computation while maintaining performance even on long inputs.

See also

Stories using this term

Browse every entry