2026
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
ICLR 2026poster
Despite its recent successes, Deep Reinforcement Learning (DRL) is notoriously sample-inefficient. We argue that this inefficiency stems from the standard practice of optimizing policies directly in the high-dimensional and highly redundant parameter space $\\Theta$. This challenge is greatly compou…