Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-world deployment conditions. Threats to the external validity of GenAI evaluations arise when the source sample of human r…