← Search

Menglin Zou

1 accepted papers

2025

Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm

NeurIPS 2025poster

Inverse Reinforcement Learning (IRL) learns a reward function to explain expert demonstrations. Modern IRL methods often use the adversarial (minimax) formulation that alternates between reward and policy optimization, which often lead to {\em unstable} training. Recent non-adversarial IRL approach…

Cited by 0SourceScholar