AI GlossaryㄷWhere everyone starts
Next Token Prediction
The way a language model builds text by repeatedly picking the single most probable next chunk to follow what's been written so far
In plain words
Think of Next Token Prediction as a much more advanced version of your smartphone keyboard's word suggestion feature. Just as the keyboard looks at what you've typed and suggests a few candidate next words, a language model looks at the text generated so far and calculates the probability of what chunk should come next, then picks one. It adds that chunk to the text, looks at the whole thing again, picks the next chunk, and repeats this process over and over until a long response is complete.
What's interesting is that this simple repeated step alone lets the model do things that seem complex, like summarizing, translating, or writing code. Correctly predicting the next chunk requires grasping not just grammar but context and common sense to some degree, so a model trained on this prediction task over vast amounts of text naturally ends up with broad capabilities. But this is also where the method's limits come from: the model isn't looking up facts to verify them, it's just selecting a chunk that plausibly continues the text, which is why it can sometimes state incorrect information with total confidence.
How it shows up in the news
Claude Academy's 'AI Capabilities and Limitations' course treats Next Token Prediction as a core concept—alongside knowledge, working memory, and context limitations—for understanding what language models can and can't do. A common misconception is dismissing it as mere autocomplete, but the course's point is that complex abilities like reasoning and summarization actually emerge from this repeated prediction process.
Try it yourself
Try giving a chatbot only half a sentence, like this:
Complete the following sentence naturally: Looking at the sky this morning...
After you get an answer, ask the model why it chose to continue that way—you'll get a feel for how it's probabilistically stringing together the next chunk based on the preceding context.
See also
Stories using this term
- Anthropic Launches Free "Academy" Site for Learning ClaudeAI · 2026.08.21
- Claude Code reveals six ways to steer hours-long tasksAI · 2026.09.04
- Claude completes first computer-verified proof of Fermat's Last TheoremAI · 2026.09.05
- Anthropic has Claude tackle AI alignment research, and it outperforms humansAI · 2026.08.31
- Claude in Chrome adds cross-device session continuityAI · 2026.08.13
- Claude tests 'Morning Brief' feature built on scheduled tasksAI · 2026.08.26
