EMNLP 20250 citations

Adaptive Platt Scaling with Causal Interpretations for Self-Reflective Language Model Uncertainty Estimates

Anthony Sicilia, Malihe Alikhani

Abstract

As large language models (LLMs) are consumed by more users and deployed in increasingly autonomous capacities, their ability to self-monitor and ask for human intervention is of vital importance. Underlying this capability are fundamental skills like self-reflection and expression of uncertainty. In this work, we provide a formal analysis of LLM self-reflection for uncertainty estimation, using domain adaptation theory to model the shift between base predictions and reflective judgments. We use this to motivate a temperature scaling algorithm that calibrates uncertainty using comparisons between base predictions and LLM self-reflections. We evaluate our approach on challenging question-answering tasks requiring reasoning, demonstrating that our methods can improve calibration of uncertainty estimates and also offer improvements in human interpretation. More broadly, this use case shows how domain adaptation presents a promising analytical tool for understanding the underlying statistical properties of LLM self-reflections.

BibTeX
@inproceedings{emnlp2025_adaptiveplattsca,
  title = {Adaptive Platt Scaling with Causal Interpretations for Self-Reflective Language Model Uncertainty Estimates},
  author = {Anthony Sicilia and Malihe Alikhani},
  booktitle = {EMNLP 2025},
  year = {2025}
}
Adaptive Platt Scaling with Causal Interpretations for Self-Reflective Language Model Uncertainty Estimates · EMNLP 2025