2018
Towards Robust Interpretability with Self-Explaining Neural Networks
NeurIPS 2018poster
Most recent work on interpretability of complex machine learning models has focused on estimating a-posteriori explanations for previously trained models around specific predictions. Self-explaining models where interpretability plays a key role already during learning have received much less attent…