One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Different Facets of Verbalised Overconfidence: an Interpretability Study

arXiv:2608.181062026-08-20

Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logical necessity and possibility, we study this behavior in Qwen3-4B, across three ways to express uncertainty: verbal epistemic markers, abstention, and numeric confidence scores. Our results confirm this tendency toward overconfidence, particularly when the model is prompted

Authors · Davide Mazzaccara, Leonardo Bertolazzi, Raffaella Bernardi

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB