매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

arXiv:2608.108122026-08-12

arXiv:2608.10812v1 Announce Type: new Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is gated by language identification. We then linearly interpolate the supervised fine-tuning (SFT) and reinforcement learning (RL) model checkpoints to ob

저자 · Chris Han, Pengzhi Gao, Pei Fu, Jian Luan

arXiv에서 원문 보기