← Search

Haocheng Xia

2 accepted papers

2026

LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional Encoding

ICML 2026poster

Key-value (KV) caching accelerates inference of large language models (LLMs) by reusing past computations for generated tokens. Its importance becomes even greater in long-context applications such as retrieval-augmented generation (RAG) and in-context learning (ICL). However, conventional KV cachin…

Cited by 0SourceScholar
2024

Data-faithful Feature Attribution: Mitigating Unobservable Confounders via Instrumental Variables

NeurIPS 2024poster

The state-of-the-art feature attribution methods often neglect the influence of unobservable confounders, posing a risk of misinterpretation, especially when it is crucial for the interpretation to remain faithful to the data. To counteract this, we propose a new approach, data-faithful feature attr…

Cited by 1SourcePDFScholar