IJCAI 2023poster6 citations

Explanation-Guided Reward Alignment

Saaduddin Mahmud, Sandhya Saisubramanian, Shlomo Zilberstein

Abstract

Agents often need to infer a reward function from observations to learn desired behaviors. However, agents may infer a reward function that does not align with the original intent because there can be multiple reward functions consistent with its observations. Operating based on such misaligned rewards can be risky. Furthermore, black-box representations make it difficult to verify the learned rewards and prevent harmful behavior. We present a framework for verifying and improving reward alignment using explanations and show how explanations can help detect misalignment and reveal failure cases in novel scenarios. The problem is formulated as inverse reinforcement learning from ranked trajectories. Verification tests created from the trajectory dataset are used to iteratively validate and improve reward alignment. The agent explains its learned reward and a tester signals whether the explanation passes the test. In cases where the explanation fails, the agent offers alternative explanations to gather feedback, which is then used to improve the learned reward. We analyze the efficiency of our approach in improving reward alignment using different types of explanations and demonstrate its effectiveness in five domains.

AI Ethics, Trust, Fairness: ETF: Safety and robustnessAI Ethics, Trust, Fairness: ETF: Explainability and interpretabilityMachine Learning: ML: Reinforcement learning
BibTeX
@inproceedings{ijcai2023p53,
  title     = {Explanation-Guided Reward Alignment},
  author    = {Mahmud, Saaduddin and Saisubramanian, Sandhya and Zilberstein, Shlomo},
  booktitle = {Proceedings of the Thirty-Second International Joint Conference on
               Artificial Intelligence, {IJCAI-23}},
  publisher = {International Joint Conferences on Artificial Intelligence Organization},
  editor    = {Edith Elkind},
  pages     = {473--482},
  year      = {2023},
  month     = {8},
  note      = {Main Track},
  doi       = {10.24963/ijcai.2023/53},
  url       = {https://doi.org/10.24963/ijcai.2023/53},
}
Explanation-Guided Reward Alignment · IJCAI 2023