AI GlossaryㅁTechnical words in the news
multi-token prediction
A technique where an AI predicts several tokens at once instead of one at a time, speeding up how fast it generates answers
In plain words
Multi-token prediction is a technique where an AI predicts several upcoming tokens (pieces of text) at once when generating a response, instead of just one at a time.
Normally, an AI generates text step by step: it produces one word, reads that back in, produces the next word, reads that back in, and so on. Think of the difference between someone typing letter by letter, versus someone who already has the next few letters in mind and types them out in one go. The second person can get more written in the same amount of time. In the same way, multi-token prediction lets an AI produce more output in the same amount of time.
The result is faster responses and lower electricity use for the same task. But because it involves guessing multiple tokens at once, there's a risk the guesses will be wrong, so in practice this is often paired with a verification step that double-checks the predictions. It's best understood as one of several techniques used to speed up AI response generation.
How it shows up in the news
Articles use it like: "Jalapeño achieved this result without using acceleration techniques such as multi-token prediction or speculative decoding." This means the system performed well even without these speed-boosting techniques, implying there's still room to get even faster by adding them later. It's also easy to misunderstand that multi-token prediction is a technology exclusive to one company — it isn't.
See also
Stories using this term
- Perplexity's local agent beats Hermes, Pi in benchmarksAI · 2026.08.26
- Google Research unveils TimesFM-3, a multivariate time-series forecasting modelAI · 2026.09.01
- Ant Group's Ling 3.0 Tiny scores 25 on intelligence index with 1.3B active parametersAI · 2026.08.12
- Gemini Cuts Video Tokens 88% With Agentic Video UnderstandingAI · 2026.09.02
- Qwen3.8-27B, run on a laptop, reportedly beats Gemini 3.7 FlashAI · 2026.08.18
- Perplexity brings cloud-free agents to DGX SparkAI · 2026.08.26
