2026
Learning from Pairwise Preferences in Long-Term Decision Problems
ICML 2026poster
Agents that can beat or tie any other under a model of pairwise preference have strong guarantees for both user satisfaction and overall social welfare. However, searching for these agents in long-term decision problems is not computationally tractable with current approaches, which require the size…