2026
Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings
ICLR 2026poster
We propose a method for evaluating the robustness of widely used LLM ranking systems---variants of a Bradley--Terry model---to dropping a worst-case very small fraction of preference data. Our approach is computationally fast and easy to adopt. When we apply our method to matchups from popular LLM r…