Post hoc Explanations may be Ineffective for Detecting Unknown Spurious Correlation
We investigate whether three types of post hoc model explanations–feature attribution, concept activation, and training point ranking–are effective for detecting a model’s reliance on spurious signals in the training data. Specifically, we consider the scenario where the spurious signal to be detect…