← Search

David Arbour

20 accepted papers

2025

Evaluation and Incident Prevention in an Enterprise AI Assistant

AAAI 2025technical

Enterprise AI Assistants are increasingly deployed in domains where accuracy is paramount, making each erroneous output a potentially significant incident. This paper presents a comprehensive framework for monitoring, benchmarking, and continuously improving such complex, multi-component systems und…

Cited by 0SourcePDFScholar
2025

Handling Missing Responses under Cluster Dependence with Applications to Language Model Evaluation

NeurIPS 2025poster

Human annotations play a crucial role in evaluating the performance of GenAI models. Two common challenges in practice, however, are missing annotations (the response variable of interest) and cluster dependence among human-AI interactions (e.g., questions asked by the same user may be highly correl…

Cited by 0SourceScholar
2025

Image Difference Captioning via Adversarial Preference Optimization

EMNLP 2025

Image Difference Captioning (IDC) aims to generate natural language descriptions that highlight subtle differences between two visually similar images. While recent advances leverage pre-trained vision-language models to align fine-grained visual differences with textual semantics, existing supervis

Cited by 0SourcePDFScholar
2025

Leveraging semantic similarity for experimentation with AI-generated treatments

NeurIPS 2025poster

Large Language Models (LLMs) enable a new form of digital experimentation where treatments combine human and model-generated content in increasingly sophisticated ways. The main methodological challenge in this setting is representing these high-dimensional treatments without losing their semantic m…

Cited by 0SourceScholar
2025

Principled Content Selection to Generate Diverse and Personalized Multi-Document Summaries

ACL 2025long

While large language models (LLMs) are increasingly capable of handling longer contexts, recent work has demonstrated that they exhibit the _”lost in the middle”_ phenomenon (Liu et al., 2024) of unevenly attending to different parts of the provided context. This hinders their ability to cover diver…

Cited by 0SourcePDFScholar
2024

Continuous Treatment Effects with Surrogate Outcomes

ICML 2024poster

In many real-world causal inference applications, the primary outcomes (labels) are often partially missing, especially if they are expensive or difficult to collect. If the missingness depends on covariates (i.e., missingness is not completely at random), analyses based on fully observed samples al…

Cited by 3SourcePDFScholar
2024

Distributional Off-Policy Evaluation for Slate Recommendations

AAAI 2024technical

Recommendation strategies are typically evaluated by using previously logged data, employing off-policy evaluation methods to estimate their expected performance. However, for strategies that present users with slates of multiple items, the resulting combinatorial action space renders many of these…

2024

Editing Partially Observable Networks via Graph Diffusion Models

ICML 2024poster

Most real-world networks are noisy and incomplete samples from an unknown target distribution. Refining them by correcting corruptions or inferring unobserved regions typically improves downstream performance. Inspired by the impressive generative capabilities that have been used to correct corrupti…

Cited by 1SourcePDFScholar
2023

Finite Population Regression Adjustment and Non-asymptotic Guarantees for Treatment Effect Estimation

NeurIPS 2023poster

The design and analysis of randomized experiments is fundamental to many areas, from the physical and social sciences to industrial settings. Regression adjustment is a popular technique to reduce the variance of estimates obtained from experiments, by utilizing information contained in auxiliary c…

Cited by 3SourcePDFScholar
2023

Learning Relational Causal Models with Cycles through Relational Acyclification

AAAI 2023technical

In real-world phenomena which involve mutual influence or causal effects between interconnected units, equilibrium states are typically represented with cycles in graphical models. An expressive class of graphical models, relational causal models, can represent and reason about complex dynamic syste…

2022

Constraint Sampling Reinforcement Learning: Incorporating Expertise for Faster Learning

AAAI 2022technical

Online reinforcement learning (RL) algorithms are often difficult to deploy in complex human-facing applications as they may learn slowly and have poor early performance. To address this, we introduce a practical algorithm for incorporating human insight to speed learning. Our algorithm, Constraint…

2022

Sample Constrained Treatment Effect Estimation

NeurIPS 2022accept

Treatment effect estimation is a fundamental problem in causal inference. We focus on designing efficient randomized controlled trials, to accurately estimate the effect of some treatment on a population of $n$ individuals. In particular, we study \textit{sample-constrained treatment effect estimati…

2020

General Identification of Dynamic Treatment Regimes Under Interference

AISTATS 2020poster

In many applied fields, researchers are ofteninterested in tailoring treatments to unit-levelcharacteristics in order to optimize an outcomeof interest. Methods for identifying andestimating treatment policies are the subjectof the dynamic treatment regime literature. Separately, in many settings th…

Cited by 13SourcePDFScholar