2020
What Did You Think Would Happen? Explaining Agent Behaviour through Intended Outcomes
NeurIPS 2020poster
We present a novel form of explanation for Reinforcement Learning, based around the notion of intended outcome. These explanations describe the outcome an agent is trying to achieve by its actions. We provide a simple proof that general methods for post-hoc explanations of this nature are impossible…