← Search

Advaith Malladi

2 accepted papers

2026

Directly Optimizing Natural Language Explanations for Behavioral Faithfulness: Simulatability and Recoverability

ICML 2026poster

Natural-language explanations are widely used to interpret machine learning models, yet many prioritize human plausibility over accurately reflecting or predicting model behavior. Prior approaches often rely on human-written rationales, producing post-hoc explanations that neither align with the mod…

Cited by 0SourceScholar
2025

Explaining Differences Between Model Pairs in Natural Language through Sample Learning

EMNLP 2025

With the growing adoption of machine learning models in critical domains, techniques for explaining differences between models have become essential for trust, debugging, and informed deployment. Previous approaches address this by identifying input transformations that cause divergent predictions o

Cited by 0SourcePDFScholar