← Search

Wei Jie Yeo

2 accepted papers

2025

Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models

EMNLP 2025

Large Language Models (LLMs) are capable of generating persuasive Natural Language Explanations (NLEs) to justify their answers. However, the faithfulness of these explanations should not be readily trusted at face value. Recent studies have proposed various methods to measure the faithfulness of NL

2025

Understanding Refusal in Language Models with Sparse Autoencoders

EMNLP 2025

Refusal is a key safety behavior in aligned language models, yet the internal mechanisms driving refusals remain opaque. In this work, we conduct a mechanistic study of refusal in instruction-tuned LLMs using sparse autoencoders to identify latent features that causally mediate refusal behaviors. We