NeurIPS 2018spotlight177 citations

Human-in-the-Loop Interpretability Prior

Isaac Lage, Andrew Ross, Samuel J Gershman, Been Kim, Finale Doshi-Velez

Abstract

We often desire our models to be interpretable as well as accurate. Prior work on optimizing models for interpretability has relied on easy-to-quantify proxies for interpretability, such as sparsity or the number of operations required. In this work, we optimize for interpretability by directly including humans in the optimization loop. We develop an algorithm that minimizes the number of user studies to find models that are both predictive and interpretable and demonstrate our approach on several data sets. Our human subjects results show trends towards different proxy notions of interpretability on different datasets, which suggests that different proxies are preferred on different tasks.

BibTeX
@inproceedings{NEURIPS2018_0a7d83f0,
 author = {Lage, Isaac and Ross, Andrew and Gershman, Samuel J and Kim, Been and Doshi-Velez, Finale},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {S. Bengio and H. Wallach and H. Larochelle and K. Grauman and N. Cesa-Bianchi and R. Garnett},
 pages = {},
 publisher = {Curran Associates, Inc.},
 title = {Human-in-the-Loop Interpretability Prior},
 url = {https://proceedings.neurips.cc/paper_files/paper/2018/file/0a7d83f084ec258aefd128569dda03d7-Paper.pdf},
 volume = {31},
 year = {2018}
}
Human-in-the-Loop Interpretability Prior · NeurIPS 2018