2026
Anytime-Valid Inference for Online Ranking of Large Language Models
ICML 2026poster
Online evaluation of large language models increasingly relies on sequentially collected pairwise preferences, enabling human-aligned assessment and continuous data collection until closely performing models can be reliably distinguished. However, adaptive sampling and continuous monitoring invalida…