2024
Ask Again, Then Fail: Large Language Models’ Vacillations in Judgment
ACL 2024long
We observe that current large language models often waver in their judgments when faced with follow-up questions, even if the original judgment was correct. This wavering presents a significant challenge for generating reliable responses and building user trust. To comprehensively assess this issue,…