2025
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
ICLR 2025poster
The release of ChatGPT in November 2022 sparked an explosion of interest in post-training and an avalanche of new preference optimization (PO) methods. These methods claim superior alignment by virtue of better correspondence with human pairwise preferences, often measured by LLM-judges. In this wor…