← Search

Zhenghao He

4 accepted papers

2026

Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders

ICLR 2026poster

Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs) by grounding outputs in retrieved evidence, but faithfulness failures, where generations contradict or extend beyond the provided sources, remain a critical challenge. Existing hallucination detection method…

Cited by 0SourcecodeScholar
2025

GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability

ICCV 2025poster

Concept Activation Vectors (CAVs) provide a powerful approach for interpreting deep neural networks by quantifying their sensitivity to human-defined concepts. However, when computed independently at different layers, CAVs often exhibit inconsistencies, making cross-layer comparisons unreliable. To…