매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

arXiv:2608.102092026-08-12

arXiv:2608.10209v1 Announce Type: new Abstract: Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. However, a key limitation of current post-training methods is the inability of human annotators and automated reward functions to faithfully capture the feedback we would like to give. We introduce Evaluation-Conditioned Training (ECT), a post-training framework that uses

저자 · Alec Harris, Kasey Corra, Archie Chaudhury, Yixiong Hao

arXiv에서 원문 보기