每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

SPADE: Self-Play in Adaptive Synthetic Executable Environments

arXiv:2608.191972026-08-18

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales. We introduce SPADE (Self-Play in Adaptive Synthetic Executable Environments), a self-play RL framework in which a single LLM plays two roles: an Environment Designer that writes complete, long-horizon training environments a

作者 · Bo Liu

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道