← Search

Daniel Nevo

2 accepted papers

2020

Investigating Gender Bias in Language Models Using Causal Mediation Analysis

NeurIPS 2020spotlight

Many interpretation methods for neural models in natural language processing investigate how information is encoded inside hidden representations. However, these methods can only measure whether the information exists, not whether it is actually used by the model. We propose a methodology grounded…