每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

arXiv:2608.172532026-08-18

Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evaluate. Self-rewarding RL reduces this dependence by enabling models to derive reward signals from their own completions. However, t

作者 · Yunhao Yang

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道