2025
Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat
ACL 2025long
Evaluating large language model (LLM) is a complex task. Pairwise ranking has emerged as state-of-the-art method to evaluate human preferences by having humans compare pairs of LLM outputs based on predefined criteria, enabling ranking across multiple LLMs by aggregating pairwise results through alg…