2023
Beyond Reward: Offline Preference-guided Policy Optimization
ICML 2023poster
This study focuses on the topic of offline preference-based reinforcement learning (PbRL), a variant of conventional reinforcement learning that dispenses with the need for online interaction or specification of reward functions. Instead, the agent is provided with fixed offline trajectories and hum…