METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅈTechnical words in the news

Direct Preference Optimization (DPO)

Direct Preference Optimization

A training method that teaches a model which of two outputs is better, rather than teaching it to hit one fixed correct answer.

In plain words

Direct Preference Optimization (DPO) retrains a model using only comparisons between outputs — judging which of two results is better — instead of pointing to a single correct answer and training the model to match it. For the same question, two answers are placed side by side, and the model learns from picks of which one feels more natural or useful.

A cooking competition judge is a useful way to picture this. Instead of scoring every dish individually, a judge can simply say "this dish beats that one," and the chef still learns which direction to improve in. DPO works the same way: it gathers a large number of these comparisons and nudges the model directly toward producing answers similar to the ones marked as better. When human comparisons aren't practical at scale, another AI model can even stand in as the judge.

This approach is especially useful for problems without one clean correct answer — for example, polishing how natural a piece of stitched-together music or writing feels. When basic training alone gets the grammar right but leaves the output feeling unpolished, an extra DPO stage can push the quality further.

How it shows up in the news

In the article, a solo developer building a piano autocomplete model used Gemini 3.5 Flash to compare pairs of completions and judge which was better, collecting preference data that was then used for DPO training. As a result, the model's outputs were picked as better than the base model's more than 69% of the time. A common misunderstanding is that DPO trains a model to be "more accurate" — it doesn't. It trains the model to move closer to the option judged better, which is a different standard from correctness.

See also

Stories using this term

Browse every entry