2026
Empirical Evidence and Analysis of a Critical Pitfall in Reward Learning from Human Feedback
IJCAI 2026
Reward learning via human feedback is a crucial capability for beneficial AI. Current methods are built on decision-making theories that assume a matched dynamics model between the learning agent and the feedback provider. However, humans often form imperfect internal dynamics models, and their feed