2025
Constructing an Optimal Behavior Basis for the Option Keyboard
NeurIPS 2025poster
Multi-task reinforcement learning aims to quickly identify solutions for new tasks with minimal or no additional interaction with the environment. Generalized Policy Improvement (GPI) addresses this by combining a set of base policies to produce a new one that is at least as good—though not necessar…