2026
Hallucination Reduction with CASAL: Contrastive Activation Steering for Amortized Learning
ICLR 2026poster
Large Language Models (LLMs) exhibit impressive capabilities but often hallucinate, confidently providing incorrect answers instead of admitting ignorance. Prior work has shown that models encode linear representations of their own knowledge and that activation steering can reduce hallucinations. Th…