매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

arXiv:2608.128512026-08-14

arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operational trajectories into executable, transferable, and inspectable procedures. Because evolution optimizes task outcomes rather than procedure safety, compromised experience can cause skill misevolution.

저자 · Xutao Mao, Liangjie Zhao, Xiang Zheng, Cong Wang

arXiv에서 원문 보기