AI GlossaryㄱWhere everyone starts
Reinforcement Learning
A trial-and-error training method that rewards good moves and penalizes bad ones. The method behind AlphaGo.
In plain words
Instead of handing over an answer key, you let the system try things, then reward it when it does well and penalize it when it doesn't. It works on the same principle as training a dog — sit on command, get a treat, repeat millions of times.
AlphaGo conquered the game of Go this way, and today it's become central to training language models. Models learn to reason better by solving math problems and getting rewarded for correct answers. "We boosted reasoning performance with reinforcement learning" has become a standard line in recent model announcements.
See also
Stories using this term
- OpenAI Publishes Safety Case Guidelines for Frontier TrainingAI · 2026.09.29
- New training method lets robot arms adjust speed like a dialAI · 2026.08.11
- Generalist AI Teaches Robots New Tasks From a 3-Second DemoAI · 2026.08.25
- Dyna Robotics unveils robot model trained on 1 million hours of human videoAI · 2026.08.11
- Unsloth releases desktop app with local training supportAI · 2026.08.12
- NVIDIA proposes robot policy trained on video instead of vision-language modelsAI · 2026.08.09
