← Search

Kimberly Truong

2 accepted papers

2026

Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas

ICLR 2026poster

As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-world deployment conditions. Threats to the external validity of GenAI evaluations arise when the source sample of human r…

Cited by 0SourceScholar
2025

Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles

EMNLP 2025

Current benchmarks for evaluating Large Language Models (LLMs) often do not exhibit enough writing style diversity, with many adhering primarily to standardized conventions. Such benchmarks do not fully capture the rich variety of communication patterns exhibited by humans. Thus, it is possible that

Cited by 0SourcePDFScholar