← Search

Justin Whitehouse

7 accepted papers

2026

Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas

ICLR 2026poster

As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-world deployment conditions. Threats to the external validity of GenAI evaluations arise when the source sample of human r…

Cited by 0SourceScholar
2024

Mutli-Armed Bandits with Network Interference

NeurIPS 2024poster

Online experimentation with interference is a common challenge in modern applications such as e-commerce and adaptive clinical trials in medicine. For example, in online marketplaces, the revenue of a good depends on discounts applied to competing goods. Statistical inference with interference is wi…

Cited by 5SourcePDFScholar
2023

Adaptive Principal Component Regression with Applications to Panel Data

NeurIPS 2023poster

Principal component regression (PCR) is a popular technique for fixed-design error-in-variables regression, a generalization of the linear regression setting in which the observed covariates are corrupted with random noise. We provide the first time-uniform finite sample guarantees for online (regul…

Cited by 7SourcePDFScholar
2022

Brownian Noise Reduction: Maximizing Privacy Subject to Accuracy Constraints

NeurIPS 2022accept

There is a disconnect between how researchers and practitioners handle privacy-utility tradeoffs. Researchers primarily operate from a privacy first perspective, setting strict privacy requirements and minimizing risk subject to these constraints. Practitioners often desire an accuracy first perspec…

Cited by 10SourcePDFScholar
2018

Efficient Formal Safety Analysis of Neural Networks

NeurIPS 2018poster

Neural networks are increasingly deployed in real-world safety-critical domains such as autonomous driving, aircraft collision avoidance, and malware detection. However, these networks have been shown to often mispredict on inputs with minor adversarial or even accidental perturbations. Consequences…