매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Progressive Content Refinement with Decaying Reward Joint LinUCB

arXiv:2608.067502026-08-10

arXiv:2608.06750v1 Announce Type: new Abstract: Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-based Self-Refine to traditional bandit approaches often rely on static options or overlook the saturation effect. This neglect leads to over-exploitation, where the continuous use of identical prompts or arms results in diminishing rewards over time. To address this challenge, we propose a novel contextual bandit

저자 · Shion Ishikawa, Pablo Loyola, Young-joo Chung, Yun Ching Liu

arXiv에서 원문 보기