Interpretable Debiasing of Vectorized Language Representations with Iterative Orthogonalization
We propose a new mechanism to augment a word vector embedding representation that offers improved bias removal while retaining the key information—resulting in improved interpretability of the representation. Rather than removing the information associated with a concept that may induce bias, our pr…