2026
Harnessing Reasoning Trajectories for Hallucination Detection via Answer-agreement Representation Shaping
ICML 2026poster
Large reasoning models (LRMs) often generate long, seemingly coherent reasoning traces yet still produce incorrect answers, making hallucination detection challenging. Although trajectories contain useful signals, directly using trace text or vanilla hidden states for detection is brittle: traces va…