One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

SPADE: Self-Play in Adaptive Synthetic Executable Environments

arXiv:2608.191972026-08-18

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales. We introduce SPADE (Self-Play in Adaptive Synthetic Executable Environments), a self-play RL framework in which a single LLM plays two roles: an Environment Designer that writes complete, long-horizon training environments a

Authors · Bo Liu

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB