METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㄷTechnical words in the news

Multi-Token Prediction head (MTP head)

A helper module that predicts several upcoming tokens at once in a single pass, speeding up generation.

In plain words

A Multi-Token Prediction head (MTP head) is an add-on that lets an AI guess several upcoming words at once, instead of writing text one word at a time in strict order.

Here's an analogy. Instead of one exam proctor solving each question and writing down the answer one by one, several assistants sit alongside and draft guesses for the next few questions in advance. The proctor glances at those drafts: if a guess is right, it's accepted and the work moves on; if it's wrong, the proctor solves it directly instead. The more often the guesses are correct, the faster the overall work goes. The part playing that assistant role is the MTP head.

This approach pays off especially in environments where simply reading stored weights is slow, such as personal computers. That's because loading the weights once can yield several tokens instead of just one. In large server environments that already batch many requests together, efficiency is already high, so the gain from an MTP head is comparatively smaller.

How it shows up in the news

In articles, it shows up in phrases like "a configuration that also includes a DSpark MTP head," meaning the model file ships with this helper module attached. A common misunderstanding is that having an MTP head automatically means several-times-faster generation — it doesn't. The actual speedup depends heavily on how often the predictions are correct and on the runtime environment, particularly whether memory bandwidth is the bottleneck.

See also

Stories using this term

Browse every entry