AI GlossaryㅋTechnical words in the news
Kullback-Leibler divergence
A numerical measure of how different two probability distributions are from each other, used to gauge how closely an AI model's predictions follow another model's.
In plain words
Kullback-Leibler divergence is a way to express, as a single number, how different two probability distributions are from each other.
Imagine two weather forecasters. One says there's an 80% chance of rain tomorrow, the other says 20%. You can sense that these two forecasts are far apart, but to actually compare them or figure out how to bring them closer together, you need to measure the gap with a number. Kullback-Leibler divergence is the mathematical yardstick that calculates exactly this kind of gap between two predictions (probability distributions). A value of 0 means the two predictions are identical; the larger the value, the further apart they are.
The same idea shows up in training AI models. Suppose a large, well-trained model (the teacher) assigns a probability to each possible answer for a given question, and a new, smaller model (the student) being trained also produces its own probabilities for the same question. If you want the student to imitate the teacher's entire probability distribution rather than just memorize the single correct answer, you calculate the difference between the two using Kullback-Leibler divergence, then nudge the student little by little to shrink that value. In other words, this metric acts as a report card showing how closely the student is coming to resemble the teacher.
How it shows up in the news
One article described shrinking a large language model to half its size and compressing it to 4-bit precision, training the student model not on correct-answer labels but purely by matching its output distribution to the teacher model's using KL divergence. A common misunderstanding here is that KL divergence isn't some new technology from a particular AI company — it's a long-established mathematical concept from probability and information theory. The article simply applied that concept to model-compression training.
Try it yourself
Giving a chatbot specific numbers like the following can help you get an intuitive feel for the concept:
"For a coin flip, I predicted a 50% chance of heads, but the actual data showed 70%. Calculate the Kullback-Leibler divergence between these two probability distributions, and explain with a simple analogy what a large value means."
See also
Stories using this term
- LiteLLM Supply Chain Attack Exposes Credentials of 2,500 OrganizationsBusiness · 2026.08.13
- DeepSeek-V4-Flash-0731 reported to stall during long-context tasksAI · 2026.08.10
- Databricks Hits $7B Run-Rate Revenue, Raises $5B Not $500MBusiness · 2026.08.14
- Apple scales up a diffusion-style language model to 1.7 billion parametersAI · 2026.08.11
- Multiverse Computing Unveils Techniques to Cut LLM Knowledge Distillation CostsAI · 2026.08.10
- China's Longsys and Enflame Both File for Hong Kong-Linked Listings Worth $800M and $900MBusiness · 2026.09.01
