2024
Calibrating Language Models with Adaptive Temperature Scaling
EMNLP 2024main
The effectiveness of large language models (LLMs) is not only measured by their ability to generate accurate outputs but also by their calibration—how well their confidence scores reflect the probability of their outputs being correct. While unsupervised pre-training has been shown to yield LLMs wit…