← Search

Ken Holstein

3 accepted papers

2026

Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas

ICLR 2026poster

As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-world deployment conditions. Threats to the external validity of GenAI evaluations arise when the source sample of human r…

Cited by 0SourceScholar
2025

Validating LLM-as-a-Judge Systems under Rating Indeterminacy

NeurIPS 2025poster

The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and standardizing GenAI evaluations. To validate such judge systems, evaluators assess human--judge agreement by first collect…

Cited by 0SourceScholar
2024

Predictive Performance Comparison of Decision Policies Under Confounding

ICML 2024poster

Predictive models are often introduced to decision-making tasks under the rationale that they improve performance over an existing decision-making policy. However, it is challenging to compare predictive performance against an existing decision-making policy that is generally under-specified and dep…