2026
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
ICLR 2026poster
Chain-of-thought explanations are widely used to inspect the decision process of large language models (LLMs) and to evaluate the trustworthiness of model outputs, making them important for effective collaboration between LLMs and humans. We demonstrate that preference optimization -- a key step in…