Beyond Drift: Stabilizing Subjective LLM Evaluation with Information-Theoretic Rubrics
Despite the growing use of large language models (LLMs) in subjective tasks such as role-playing, humor, emotional intelligence, and dialogue quality, their evaluation faces a pressing reproducibility crisis: even the same evaluator may contradict itself when re-judging the exact same sample. We att…