매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

arXiv:2608.116942026-08-13

arXiv:2608.11694v1 Announce Type: new Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions, so some failures become successes and some successes become failures. We call this drift. BenchDrift generates meaning-preserving var

저자 · Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel

arXiv에서 원문 보기