2026
Correct looks better: Pairwise comparisons reveal accuracy rankings
ICML 2026poster
Pairwise comparisons by humans or judge models, combined with aggregation methods such as Elo or Bradley-Terry, have become a central part of evaluating generative models. However, there has been significant debate whether they measure what they intend to measure. Some argue, pairwise comparisons fr…