2022
UNIREX: A Unified Learning Framework for Language Model Rationale Extraction
ICML 2022spotlight
An extractive rationale explains a language model’s (LM’s) prediction on a given task instance by highlighting the text inputs that most influenced the prediction. Ideally, rationale extraction should be faithful (reflective of LM’s actual behavior) and plausible (convincing to humans), without comp…