Confidence is Not Universal: Task-Dependent Calibration and Emergent Behavior in LLMs
Large language models (LLMs) increasingly support human decision-making, rendering human-interpretable confidence essential. However, it remains unclear whether verbalized confidence calibration generalizes across heterogeneous tasks without degrading accuracy. We show that universal confidence cali…