← Search

Harold Abelson

1 accepted papers

2022

Post hoc Explanations may be Ineffective for Detecting Unknown Spurious Correlation

ICLR 2022poster

We investigate whether three types of post hoc model explanations–feature attribution, concept activation, and training point ranking–are effective for detecting a model’s reliance on spurious signals in the training data. Specifically, we consider the scenario where the spurious signal to be detect…

Cited by 118SourcePDFScholar