← Search

Mengyu Ye

3 accepted papers

2025

Can Input Attributions Explain Inductive Reasoning in In-Context Learning?

ACL 2025finding

Interpreting the internal process of neural models has long been a challenge. This challenge remains relevant in the era of large language models (LLMs) and in-context learning (ICL); for example, ICL poses a new issue of interpreting which example in the few-shot examples contributed to identifying…

Cited by 0SourcePDFScholar
2025

Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders

NeurIPS 2025poster

Recent interpretability work on large language models (LLMs) has been increasingly dominated by a feature-discovery approach with the help of proxy modules. Then, the quality of features learned by, e.g., sparse auto-encoders (SAEs), is evaluated. This paradigm naturally raises a critical question:…

Cited by 0SourceScholar
2023

Assessing Step-by-Step Reasoning against Lexical Negation: A Case Study on Syllogism

EMNLP 2023short main

Large language models (LLMs) take advantage of step-by-step reasoning instructions, e.g., chain-of-thought (CoT) prompting. Building on this, their ability to perform CoT-style reasoning robustly is of interest from a probing perspective. In this study, we inspect the step-by-step reasoning abilit…

Cited by 0SourceScholar