← Search

Sarah Dean

18 accepted papers

2026

Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking

ICML 2026poster

Large-scale search, recommendation, and retrieval-augmented generation (RAG) systems typically employ a two-stage architecture: an early-stage ranker (ESR) generates a candidate set, which is subsequently re-ranked by a late-stage ranker (LSR). While there are many reinforcement learning (RL) method…

Cited by 0SourceScholar
2026

High-Altitude Balloon Station-Keeping with First Order Model Predictive Control

ICRA 2026poster

High-altitude balloons (HABs) are common in scientific research due to their wide range of applications and low cost. Because of their nonlinear, underactuated dynamics and the partial observability of wind fields, prior work has largely relied on model-free reinforcement learning (RL) methods to de…

2026

Two-Layer Linear Auto-Regressive Models Estimate Latent States

ICML 2026poster

Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models learn latent representations remains an open theoretical question. In this work, we demonstrate that when trained by empirical risk minimization on data from part…

Cited by 0SourceScholar
2025

Policy Design for Two-sided Platforms with Participation Dynamics

ICML 2025poster

In two-sided platforms (e.g., video streaming or e-commerce), viewers and providers engage in interactive dynamics: viewers benefit from increases in provider populations, while providers benefit from increases in viewer population. Despite the importance of such “population effects” on long-term pl…

2025

Pre-trained Large Language Models Learn to Predict Hidden Markov Models In-context

NeurIPS 2025poster

Hidden Markov Models (HMMs) are fundamental tools for modeling sequential data with latent states that follow Markovian dynamics. However, they present significant challenges in model fitting and computational efficiency on real-world datasets. In this work, we demonstrate that pre-trained large l…

Cited by 0SourceScholar
2025

To Ask or not to Ask: Human-in-the-loop Contextual Bandits with Applications in Robot-Assisted Feeding

ICRA 2025

Robot-assisted bite acquisition involves picking up food items with varying shapes, compliance, sizes, and textures. Fully autonomous strategies may not generalize efficiently across this diversity. We propose leveraging feedback from the care recipient when encountering novel food items. However, f

Cited by 9SourceScholar
2024

Emergent specialization from participation dynamics and multi-learner retraining

AISTATS 2024poster

Numerous online services are data-driven: the behavior of users affects the system’s parameters, and the system’s parameters affect the users’ experience of the service, which in turn affects the way users may interact with the system. For example, people may choose to use a service only for tasks t…

2024

Initializing Services in Interactive ML Systems for Diverse Users

NeurIPS 2024poster

This paper investigates ML systems serving a group of users, with multiple models/services, each aimed at specializing to a sub-group of users. We consider settings where upon deploying a set of services, users choose the one minimizing their personal losses and the learner iteratively learns by int…

Cited by 10SourcePDFScholar
2023

Modeling content creator incentives on algorithm-curated platforms

ICLR 2023top-5%

Content creators compete for user attention. Their reach crucially depends on algorithmic choices made by developers on online platforms. To maximize exposure, many creators adapt strategically, as evidenced by examples like the sprawling search engine optimization industry. This begets competition…

Cited by 44SourcePDFScholar
2021

Quantifying Availability and Discovery in Recommender Systems via Stochastic Reachability

ICML 2021spotlight

In this work, we consider how preference models in interactive recommendation systems determine the availability of content and users’ opportunities for discovery. We propose an evaluation procedure based on stochastic reachability to quantify the maximum probability of recommending a target piece o…

2020

Balancing Competing Objectives with Noisy Data: Score-Based Classifiers for Welfare-Aware Machine Learning

ICML 2020poster

While real-world decisions involve many competing objectives, algorithmic decisions are often evaluated with a single objective function. In this paper, we study algorithmic policies which explicitly trade off between a private objective (such as profit) and a public objective (such as social welfar…

2020

Guaranteeing Safety of Learned Perception Modules via Measurement-Robust Control Barrier Functions

CoRL 2020

Modern nonlinear control theory seeks to develop feedback controllers that endow systems with properties such as safety and stability. The guarantees ensured by these controllers often rely on accurate estimates of the system state for determining control actions. In practice, measurement model unce

2018

Delayed Impact of Fair Machine Learning

ICML 2018oral

Fairness in machine learning has predominantly been studied in static classification settings without concern for how decisions change the underlying population over time. Conventional wisdom suggests that fairness criteria promote the long-term well-being of those groups they aim to protect. We stu…

2018

Regret Bounds for Robust Adaptive Control of the Linear Quadratic Regulator

NeurIPS 2018poster

We consider adaptive control of the Linear Quadratic Regulator (LQR), where an unknown linear system is controlled subject to quadratic costs. Leveraging recent developments in the estimation of linear systems and in robust controller synthesis, we present the first provably polynomial time algorith…

Cited by 332SourcePDFScholar