← Search

Katie Matton

1 accepted papers

2025

Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations

ICLR 2025spotlight

Large language models (LLMs) are capable of generating *plausible* explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can be *unfaithful*. This, in turn, can lead to over-trust and misuse. We introduce…