2026
A Positive Case for Faithfulness: Explanations Help Predict Model Behavior
ICML 2026poster
LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithfulness metrics have critical limitations, typically relying on identifying unfaithfulness via adversarial prompting or det…