METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㄱTechnical words in the news

cross-entropy loss

A training score that measures how far a model's probability predictions are from the correct answer

In plain words

Cross-entropy loss is a numeric grading method used during training to score how far off an AI model's answer was from the correct one. Think of a multiple-choice test grading scheme: if a student confidently picks the one right answer, they get near full marks, but the more confidently they pick a wrong answer, the bigger the penalty. When a model picks the next word or note, it spreads probability across several candidates — if it gave the actual correct answer a low probability, the penalty is large; if it gave it a high probability, the penalty is small. Training then nudges the model little by little to shrink that penalty.

The catch is that this grading is short-sighted, judging only the 'very next step.' It can build a habit of nailing the right answer at each individual moment, but whether the whole finished output — built from many such steps — sounds natural and coherent overall is a separate question. That's why cross-entropy is widely used for early-stage training that teaches basic grammar, but later stages that polish the overall quality of a finished output often add other methods on top.

How it shows up in the news

Regarding a piano autocomplete model, the article explains that "the cross-entropy approach, which only treats the single next note in a held-out song as the 'correct answer' during training, teaches the basic grammar of music but has limits when it comes to actually producing full pieces that sound natural." A common misunderstanding is that a low cross-entropy score doesn't necessarily mean the finished output sounds good. In this article too, the developer followed up cross-entropy training with an additional round of training that compared two outputs and picked the better one.

Try it yourself

Ask a chatbot this to get the concept: "Explain cross-entropy loss using a multiple-choice test grading analogy, simple enough for an elementary school student. Also explain how the score differs when the model gives a low probability versus a high probability to the correct answer."

See also

Stories using this term

Browse every entry