← Search

Gael Gendron

2 accepted papers

2025

Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm

NeurIPS 2025poster

Inverse Reinforcement Learning (IRL) learns a reward function to explain expert demonstrations. Modern IRL methods often use the adversarial (minimax) formulation that alternates between reward and policy optimization, which often lead to {\em unstable} training. Recent non-adversarial IRL approach…

Cited by 0SourceScholar
2024

Can Large Language Models Learn Independent Causal Mechanisms?

EMNLP 2024main

Despite impressive performance on language modelling and complex reasoning tasks, Large Language Models (LLMs) fall short on the same tasks in uncommon settings or with distribution shifts, exhibiting a lack of generalisation ability. By contrast, systems such as causal models, that learn abstract v…