Better Than Diverse Demonstrators: Reward Decomposition From Suboptimal and Heterogeneous Demonstrations
Inverse Reinforcement Learning (IRL) typically involves inferring a reward function from expert demonstrations to enable agents to imitate the demonstrated behavior. However, real-world settings often provide suboptimal and heterogeneous demonstrations, where human demonstrators use diverse strategi