TREND: Tri-Teaching for Robust Preference-based Reinforcement Learning with Demonstrations
Preference feedback collected by human or VLM annotators is often noisy, presenting a significant challenge for preference-based reinforcement learning that relies on accurate preference labels. To address this challenge, we propose TREND, a novel framework that integrates few-shot expert demonstrat