2025
Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector
NeurIPS 2025poster
LLM-as-a-Judge has emerged as a promising tool for automatically evaluating generated outputs, but its reliability is often undermined by potential biases in judgment. Existing efforts to mitigate these biases face key limitations: in-context learning-based methods fail to address rooted biases due…