매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

arXiv:2608.075282026-08-11

arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured confidence formats collapse to two values with indistinguishable e

저자 · Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk

arXiv에서 원문 보기