2026
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
ICML 2026poster
While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly when tasks demand information outside their parametric knowledge. Indeed, truthfulness requires more than accuracy---mod…