Efficient and Near-Optimal Algorithm for Contextual Dueling Bandits with Offline Regression Oracles
The problem of contextual dueling bandits is central to reinforcement learning with human feedback (RLHF), a widely used approach in AI alignment for incorporating human preferences into learning systems. Despite its importance, existing methods are constrained either by strong preference modeling a…