2026
Rethinking the Hardness of PbRL: A Provable General Regret Bound
ICML 2026poster
This paper studies \emph{preference-based reinforcement learning} (PbRL), where agents learn from comparative, trajectory-level feedback rather than numeric rewards. While PbRL has seen rapid empirical and theoretical progress, existing analyses are largely confined to restricted settings and fail t…