← Search

Xiangyu Hong

3 accepted papers

2025

DePass: Unified Feature Attributing by Simple Decomposed Forward Pass

NeurIPS 2025poster

Attributing the behavior of Transformer models to internal computations is a central challenge in mechanistic interpretability. We introduce DePass, a unified framework for feature attribution based on a single decomposed forward pass. DePass decomposes hidden states into customized additive compone…

Cited by 0SourceScholar
2024

On Large Language Models’ Hallucination with Regard to Known Facts

NAACL 2024long

Large language models are successful in answering factoid questions but are also prone to hallucination.We investigate the phenomenon of LLMs possessing correct answer knowledge yet still hallucinating from the perspective of inference dynamics, an area not previously covered in studies on hallucina…

2024

On the token distance modeling ability of higher RoPE attention dimension

EMNLP 2024finding

Length extrapolation algorithms based on Rotary position embedding (RoPE) have shown promising results in extending the context length of language models. However, understanding how position embedding can capture longer-range contextual information remains elusive. Based on the intuition that differ…

Cited by 5SourcePDFScholar