2025
Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models
EMNLP 2025
Large Language Models (LLMs) are capable of generating persuasive Natural Language Explanations (NLEs) to justify their answers. However, the faithfulness of these explanations should not be readily trusted at face value. Recent studies have proposed various methods to measure the faithfulness of NL