2025
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
ACL 2025long
Language models are often miscalibrated, leading to confidently incorrect answers. We introduce GRACE, a benchmark for language model calibration that incorporates comparison with human calibration. GRACE consists of question-answer pairs, in which each question contains a series of clues that gradu…