ICLR 2024poster2 citations

Identifying Policy Gradient Subspaces

Jan Schneider, Pierre Schumacher, Simon Guist, Le Chen, Daniel Haeufle, Bernhard Schölkopf, Dieter Büchler

Abstract

Policy gradient methods hold great potential for solving complex continuous control tasks. Still, their training efficiency can be improved by exploiting structure within the optimization problem. Recent work indicates that supervised learning can be accelerated by leveraging the fact that gradients lie in a low-dimensional and slowly-changing subspace. In this paper, we conduct a thorough evaluation of this phenomenon for two popular deep policy gradient methods on various simulated benchmark tasks. Our results demonstrate the existence of such gradient subspaces despite the continuously changing data distribution inherent to reinforcement learning. These findings reveal promising directions for future work on more efficient reinforcement learning, e.g., through improving parameter-space exploration or enabling second-order optimization.

reinforcement learningpolicy gradientsgradient subspaces
BibTeX
@inproceedings{
schneider2024identifying,
title={Identifying Policy Gradient Subspaces},
author={Jan Schneider and Pierre Schumacher and Simon Guist and Le Chen and Daniel Haeufle and Bernhard Sch{\"o}lkopf and Dieter B{\"u}chler},
booktitle={The Twelfth International Conference on Learning Representations},
year={2024},
url={https://openreview.net/forum?id=iPWxqnt2ke}
}
Identifying Policy Gradient Subspaces · ICLR 2024