AI GlossaryㅇTechnical words in the news
Agent RL
A training method where a model actually uses tools and improves by being rewarded for success or failure
In plain words
Agent RL is a training approach that teaches an AI model to use tools by actually doing the work, not just reading about it. Learning to cook by memorizing a recipe is different from standing at the stove, cutting ingredients, and adjusting the heat yourself. At first you might burn something, but you get better through trial and error, praised for success and correcting after failure. Agent RL works the same way: the model is given real tasks like fixing code, running terminal commands, or searching the web, and its results are scored. That score is then used to gradually fine-tune the model.
This training usually happens in stages. It often starts with tasks that have clear right answers, like math problems, then expands into software tasks, terminal use, and web search, increasing in difficulty and scope along the way. It's similar to a relay race, where the output of one stage becomes the starting point for training the next.
Because this kind of training is costly and complex, it isn't applied equally across all model sizes. Smaller models are often trained only up to following instructions, while larger models get the additional training needed to actually operate tools.
How it shows up in the news
When introducing Granite 4.2, IBM said the 8B and 30B models included an Agent RL block progressing through software engineering, terminal use, and web search, while the 3B model skipped this stage. The easy mistake here is assuming Agent RL is a single process applied uniformly to a model — it's actually a separate training stage that can be selectively included or skipped depending on model size. Even models sharing the same name can differ in whether they've actually learned to use tools, depending on their size.
See also
Stories using this term
- AWS Adds Open-Source Agent Skills for Bedrock Automated Reasoning PoliciesAI · 2026.08.09
- Google Research unveils AgentHands, an XR agent that gesturesAI · 2026.08.26
- AWS unveils bridge letting cloud agents use local MCP toolsAI · 2026.08.09
- Prime Intellect unveils self-improving agent harness 'Prime Agent'AI · 2026.08.09
- AWS adds AI traffic rate limiting to AgentCore gatewayAI · 2026.08.09
- OpenAI tightens monitoring and isolation after Hugging Face incidentBusiness · 2026.08.19
