2025
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
ICLR 2025oral
Hallucinations in large language models are a widespread problem, yet the mechanisms behind whether models will hallucinate are poorly understood, limiting our ability to solve this problem. Using sparse autoencoders as an interpretability tool, we discover that a key part of these mechanisms is ent…