METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅁTechnical words in the news

multi-token prediction

A technique where an AI predicts several tokens at once instead of one at a time, speeding up how fast it generates answers

In plain words

Multi-token prediction is a technique where an AI predicts several upcoming tokens (pieces of text) at once when generating a response, instead of just one at a time.

Normally, an AI generates text step by step: it produces one word, reads that back in, produces the next word, reads that back in, and so on. Think of the difference between someone typing letter by letter, versus someone who already has the next few letters in mind and types them out in one go. The second person can get more written in the same amount of time. In the same way, multi-token prediction lets an AI produce more output in the same amount of time.

The result is faster responses and lower electricity use for the same task. But because it involves guessing multiple tokens at once, there's a risk the guesses will be wrong, so in practice this is often paired with a verification step that double-checks the predictions. It's best understood as one of several techniques used to speed up AI response generation.

How it shows up in the news

Articles use it like: "Jalapeño achieved this result without using acceleration techniques such as multi-token prediction or speculative decoding." This means the system performed well even without these speed-boosting techniques, implying there's still room to get even faster by adding them later. It's also easy to misunderstand that multi-token prediction is a technology exclusive to one company — it isn't.

See also

Stories using this term

Browse every entry