OPRIDE: Efficient Offline Preference-based Reinforcement Learning via In-Dataset Exploration
Preference-based reinforcement learning (PbRL) can help avoid sophisticated reward designs and align better with human intentions, showing great promise in various real-world applications. However, obtaining human feedback for preferences can be expensive and time-consuming, which forms a strong bar…