매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

SPADE: Self-Play in Adaptive Synthetic Executable Environments

arXiv:2608.191972026-08-18

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales. We introduce SPADE (Self-Play in Adaptive Synthetic Executable Environments), a self-play RL framework in which a single LLM plays two roles: an Environment Designer that writes complete, long-horizon training environments a

저자 · Bo Liu

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사