← Search

Usman Anwar

3 accepted papers

2025

Interpreting Emergent Planning in Model-Free Reinforcement Learning

ICLR 2025oral

We present the first mechanistic evidence that model-free reinforcement learning agents can learn to plan. This is achieved by applying a methodology based on concept-based interpretability to a model-free agent in Sokoban -- a commonly used benchmark for studying planning. Specifically, we demonstr…

Cited by 0SourcePDFScholar
2024

Reward Model Ensembles Help Mitigate Overoptimization

ICLR 2024poster

Reinforcement learning from human feedback (RLHF) is a standard approach for fine-tuning large language models to follow instructions. As part of this process, learned reward models are used to approximately model human preferences. However, as imperfect representations of the “true” reward, these l…