2025
Firm or Fickle? Evaluating Large Language Models Consistency in Sequential Interactions
ACL 2025finding
Large Language Models (LLMs) have shown remarkable capabilities across various tasks, but their deployment in high-stake domains requires consistent and coherent behavior across multiple rounds of user interaction. This paper introduces a comprehensive framework for evaluating and improving LLM resp…