2026
Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges
ICLR 2026poster
Reliable certification of Large Language Models (LLMs)—verifying that failure rates are below a safety threshold—is critical yet challenging. While "LLM-as-a-Judge" offers scalability, judge imperfections, noise, and bias can invalidate statistical guarantees. We introduce a "Noisy but Valid" hypoth…