Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
arXiv:2608.08283v1 Announce Type: new Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error t
arXiv에서 원문 보기