2026
Improving Zero-Shot Offline RL via Behavioral Task Sampling
ICML 2026poster
Offline zero-shot reinforcement learning (RL) aims to learn agents that optimize unseen reward functions without additional environment interaction. The standard approach to this problem trains task-conditioned policies by sampling task vectors that define linear reward functions over learned state …