← Search

Lily Xu

13 accepted papers

2026

Generative AI Against Poaching: Latent Composite Flow Matching for Poaching Prediction

AAAI 2026technical

Poaching poses significant threats to biodiversity. A valuable step in reducing poaching is to forecast poacher behavior, which can inform patrol deployment and other conservation interventions. Existing poaching prediction methods based on linear models or decision trees lack the expressivity to ca

Cited by 0SourcePDFScholar
2026

Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions

ICML 2026spotlight

Reinforcement learning (RL) with combinatorial action spaces remains challenging because feasible action sets are exponentially large and governed by complex feasibility constraints, making direct policy parameterization impractical. Existing approaches embed task-specific value functions into const…

Cited by 0SourceScholar
2025

Context in Public Health for Underserved Communities: A Bayesian Approach to Online Restless Bandits

AAAI 2025technical

Public health programs often provide interventions to encourage program adherence, and effectively allocating interventions is vital for producing the greatest overall health outcomes, especially in underserved communities where resources are limited. Such resource allocation problems are often mode…

2025

Reinforcement learning with combinatorial actions for coupled restless bandits

ICLR 2025poster

Reinforcement learning (RL) has increasingly been applied to solve real-world planning problems, with progress in handling large state spaces and time horizons. However, a key bottleneck in many domains is that RL methods cannot accommodate large, combinatorially structured action spaces. In such se…

2023

Flexible Budgets in Restless Bandits: A Primal-Dual Algorithm for Efficient Budget Allocation

AAAI 2023technical

Restless multi-armed bandits (RMABs) are an important model to optimize allocation of limited resources in sequential decision-making settings. Typical RMABs assume the budget --- the number of arms pulled --- to be fixed for each step in the planning horizon. However, for realistic real-world plann…

Cited by 8SourcePDFScholar
2023

Optimistic Whittle Index Policy: Online Learning for Restless Bandits

AAAI 2023technical

Restless multi-armed bandits (RMABs) extend multi-armed bandits to allow for stateful arms, where the state of each arm evolves restlessly with different transitions depending on whether that arm is pulled. Solving RMABs requires information on transition dynamics, which are often unknown upfront. T…

2023

Robust Planning over Restless Groups: Engagement Interventions for a Large-Scale Maternal Telehealth Program

AAAI 2023technical

In 2020, maternal mortality in India was estimated to be as high as 130 deaths per 100K live births, nearly twice the UN's target. To improve health outcomes, the non-profit ARMMAN sends automated voice messages to expecting and new mothers across India. However, 38% of mothers stop listening to the…

Cited by 10SourcePDFScholar
2022

Coordinating Followers to Reach Better Equilibria: End-to-End Gradient Descent for Stackelberg Games

AAAI 2022technical

A growing body of work in game theory extends the traditional Stackelberg game to settings with one leader and multiple followers who play a Nash equilibrium. Standard approaches for computing equilibria in these games reformulate the followers' best response as constraints in the leader's optimizat…

Cited by 31SourcePDFScholar
2022

Ranked Prioritization of Groups in Combinatorial Bandit Allocation

IJCAI 2022poster

Preventing poaching through ranger patrols protects endangered wildlife, directly contributing to the UN Sustainable Development Goal 15 of life on land. Combinatorial bandits have been used to allocate limited patrol resources, but existing approaches overlook the fact that each location is home to…

2022

Restless and uncertain: Robust policies for restless bandits via deep multi-agent reinforcement learning

UAI 2022poster

We introduce robustness in \textit{restless multi-armed bandits} (RMABs), a popular model for constrained resource allocation among independent stochastic processes (arms). Nearly all RMAB techniques assume stochastic dynamics are precisely known. However, in many real-world settings, dynamics are e…

2021

Robust reinforcement learning under minimax regret for green security

UAI 2021poster

Green security domains feature defenders who plan patrols in the face of uncertainty about the adversarial behavior of poachers, illegal loggers, and illegal fishers. Importantly, the deterrence effect of patrols on adversaries’ future behavior makes patrol planning a sequential decision-making prob…