2025
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
EMNLP 2025
Large language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, often resulting in inconsistent or unreliable generated text. Different methods have been proposed to mitigate such hallucinations and fragility, one of which is to measure the consistency of LLM response