2021
Contrastive Explanations for Reinforcement Learning via Embedded Self Predictions
ICLR 2021oral
We investigate a deep reinforcement learning (RL) architecture that supports explaining why a learned agent prefers one action over another. The key idea is to learn action-values that are directly represented via human-understandable properties of expected futures. This is realized via the embedded…