← Search

Anthony GX-Chen

3 accepted papers

2026

KL-Regularized Reinforcement Learning is Designed to Mode Collapse

ICLR 2026poster

Classical intuitions cast minimizing reverse KL as "mode seeking" and forward KL as "mass covering". In KL-regularized reinforcement learning, however, the regularizer determines _both_ the target distribution's shape _and_ the divergence being implicitly minimized, making its role more nuanced than…

Cited by 0SourceScholar
2025

Efficient Exploration and Discriminative World Model Learning with an Object-Centric Abstraction

ICLR 2025poster

In the face of difficult exploration problems in reinforcement learning, we study whether giving an agent an object-centric mapping (describing a set of items and their attributes) allow for more efficient learning. We found this problem is best solved hierarchically by modelling items at a higher l…

Cited by 0SourcePDFScholar
2022

A Generalized Bootstrap Target for Value-Learning, Efficiently Combining Value and Feature Predictions

AAAI 2022technical

Estimating value functions is a core component of reinforcement learning algorithms. Temporal difference (TD) learning algorithms use bootstrapping, i.e. they update the value function toward a learning target using value estimates at subsequent time-steps. Alternatively, the value function can be u…

Cited by 1SourcePDFScholar