2025
Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm
NeurIPS 2025poster
Inverse Reinforcement Learning (IRL) learns a reward function to explain expert demonstrations. Modern IRL methods often use the adversarial (minimax) formulation that alternates between reward and policy optimization, which often lead to {\em unstable} training. Recent non-adversarial IRL approach…