2026
Directly Optimizing Natural Language Explanations for Behavioral Faithfulness: Simulatability and Recoverability
ICML 2026poster
Natural-language explanations are widely used to interpret machine learning models, yet many prioritize human plausibility over accurately reflecting or predicting model behavior. Prior approaches often rely on human-written rationales, producing post-hoc explanations that neither align with the mod…