2025
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
ACL 2025finding
Psychology research has shown that humans are poor at estimating their performance on tasks, tending towards underconfidence on easy tasks and overconfidence on difficult tasks. We examine three LLMs, Llama-3-70B-instruct, Claude-3-Sonnet, and GPT-4o, on a range of QA tasks of varying difficulty, an…