每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

Different Facets of Verbalised Overconfidence: an Interpretability Study

arXiv:2608.181062026-08-20

Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logical necessity and possibility, we study this behavior in Qwen3-4B, across three ways to express uncertainty: verbal epistemic markers, abstention, and numeric confidence scores. Our results confirm this tendency toward overconfidence, particularly when the model is prompted

作者 · Davide Mazzaccara, Leonardo Bertolazzi, Raffaella Bernardi

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道