2025
Task Calibration: Calibrating Large Language Models on Inference Tasks
ACL 2025finding
Large language models (LLMs) have exhibited impressive zero-shot performance on inference tasks. However, LLMs may suffer from spurious correlations between input texts and output labels, which limits LLMs’ ability to reason based purely on general language understanding. For example, in the natural…