METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅌTechnical words in the news

Tree-RL

A reinforcement learning training method that branches into multiple paths mid-run and learns by trying each of them out

In plain words

Tree-RL is a training method where an AI learns better decisions through trial and error, but instead of running a single attempt from start to finish along one path, it creates branch points partway through and tries out multiple paths at once from those points.

Think of it like playing a game but instead of playing through once and stopping, you set up save points at key moments and replay from those points with different choices multiple times. Just as a tree splits from its trunk into many branches, starting from the same point and branching in different directions lets you compare the outcomes of each branch, making it far more efficient to learn which choices were actually better.

For this to actually work, the AI needs technology that can rewind and clone the state it was working in back to a specific point in time. Without that, there's no way to create a branch point, since the run can only move forward.

How it shows up in the news

The article reports that "in Tree-RL training, which branches rollouts from a chosen point, the TerminalBench-2 score rose from 34.2% to 39.4%." It's worth noting that Tree-RL isn't a specific company's product — it's a training technique that only becomes possible with a runtime (Shepherd) capable of rewinding and cloning execution state.

See also

Stories using this term

Browse every entry