每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

arXiv:2608.173932026-08-17

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train-inference discrepancies decouple rollout behavior from policy updates. To address this, we present LEGO-RL, a framework that bridges nati

作者 · Yiming Du

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道