AI GlossaryㅈTechnical words in the news
Direct Preference Optimization (DPO)
Direct Preference Optimization
A training method that teaches a model which of two outputs is better, rather than teaching it to hit one fixed correct answer.
In plain words
Direct Preference Optimization (DPO) retrains a model using only comparisons between outputs — judging which of two results is better — instead of pointing to a single correct answer and training the model to match it. For the same question, two answers are placed side by side, and the model learns from picks of which one feels more natural or useful.
A cooking competition judge is a useful way to picture this. Instead of scoring every dish individually, a judge can simply say "this dish beats that one," and the chef still learns which direction to improve in. DPO works the same way: it gathers a large number of these comparisons and nudges the model directly toward producing answers similar to the ones marked as better. When human comparisons aren't practical at scale, another AI model can even stand in as the judge.
This approach is especially useful for problems without one clean correct answer — for example, polishing how natural a piece of stitched-together music or writing feels. When basic training alone gets the grammar right but leaves the output feeling unpolished, an extra DPO stage can push the quality further.
How it shows up in the news
In the article, a solo developer building a piano autocomplete model used Gemini 3.5 Flash to compare pairs of completions and judge which was better, collecting preference data that was then used for DPO training. As a result, the model's outputs were picked as better than the base model's more than 69% of the time. A common misunderstanding is that DPO trains a model to be "more accurate" — it doesn't. It trains the model to move closer to the option judged better, which is a different standard from correctness.
See also
Stories using this term
- Tencent open-sources Hy4 preview, edges out GPT-5.6 on coding benchmarkAI · 2026.08.30
- Apple Unveils BDHS, an Alignment Technique to Reduce Multimodal AI HallucinationAI · 2026.08.11
- Gilbert+Tobin automates legal operations with 87% ChatGPT active-use rateAI · 2026.09.02
- PDI Unveils Agentic Internal App Deployment System Built on AWSAI · 2026.08.09
- Claude Beats Industry Average Success Rate in Designing Drug-Binding ProteinsAI · 2026.08.19
- OpenAI tightens monitoring and isolation after Hugging Face incidentBusiness · 2026.08.19
