2024
Zeroth-Order Optimization Meets Human Feedback: Provable Learning via Ranking Oracles
ICLR 2024poster
In this study, we delve into an emerging optimization challenge involving a black-box objective function that can only be gauged via a ranking oracle—a situation frequently encountered in real-world scenarios, especially when the function is evaluated by human judges. A prominent instance of such a…