← Search

Milind Tambe

57 accepted papers

2026

Adaptive Multi-Round Allocation with Stochastic Arrivals

ICML 2026poster

We study a sequential resource allocation problem motivated by adaptive network recruitment, in which a limited budget of identical resources must be allocated over multiple rounds to individuals with stochastic referral capacity. Successful referrals endogenously generate future decision opportunit…

Cited by 0SourceScholar
2026

Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information

ICML 2026poster

With the rapid progress of multi-agent large language model (LLM) reasoning, how to effectively aggregate answers from multiple LLMs has emerged as a fundamental challenge. Standard majority voting treats all answers equally, failing to consider latent heterogeneity and correlation across models. In…

Cited by 0SourceScholar
2026

Generative AI Against Poaching: Latent Composite Flow Matching for Poaching Prediction

AAAI 2026technical

Poaching poses significant threats to biodiversity. A valuable step in reducing poaching is to forecast poacher behavior, which can inform patrol deployment and other conservation interventions. Existing poaching prediction methods based on linear models or decision trees lack the expressivity to ca

Cited by 0SourcePDFScholar
2026

Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions

ICML 2026spotlight

Reinforcement learning (RL) with combinatorial action spaces remains challenging because feasible action sets are exponentially large and governed by complex feasibility constraints, making direct policy parameterization impractical. Existing approaches embed task-specific value functions into const…

Cited by 0SourceScholar
2026

Policy-Embedded Graph Expansion: Networked HIV Testing with Diffusion-Driven Network Samples

IJCAI 2026

HIV is a retrovirus that attacks the human immune system and can lead to death without proper treatment. In collaboration with the WHO and a large South African university, we study how to improve the efficiency of HIV testing with the goal of eventual deployment, directly supporting progress toward

Cited by 0Scholar
2026

Preference Robustness for DPO with Applications to Public Health

AAAI 2026technical

We study an LLM fine-tuning task for designing reward functions for sequential resource allocation problems in public health, guided by human preferences expressed in natural language. This setting presents a challenging testbed for alignment due to complex and ambiguous objectives and limited data

Cited by 0SourcePDFScholar
2026

Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective

ICML 2026poster

Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy. This practice is suboptimal for maximizing user's utility because the KL regularization may cause the LLM to inherit the…

Cited by 0SourceScholar
2026

Rule-Bottleneck RL: Learning to Decide and Explain for Sequential Resource Allocation via LLM Agents in Public Health

IJCAI 2026

Reducing preventable maternal mortality remains a global health priority. Under Sustainable Development Goal (SDG) target 3.1, the WHO emphasizes timely and equitable allocation of limited maternal health resources. Motivated by Department of Obstetrics and Gynecology at several important hospitals

Cited by 0Scholar
2026

Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations

ICLR 2026poster

As humans increasingly share environments with diverse agents powered by RL, LLMs, and beyond, the ability to explain agent policies in natural language is vital for reliable coexistence. We introduce a general-purpose framework that trains explanation-generating LLMs via reinforcement learning from…

Cited by 0SourceScholar
2025

Adaptive Frontier Exploration on Graphs with Applications to Network-Based Disease Testing

NeurIPS 2025poster

We study a sequential decision-making problem on a $n$-node graph $\mathcal{G}$ where each node has an unknown label from a finite set $\mathbf{\Omega}$, drawn from a joint distribution $\mathcal{P}$ that is Markov with respect to $\mathcal{G}$. At each step, selecting a node reveals its label and y…

Cited by 0SourceScholar
2025

Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data

NeurIPS 2025spotlight

Incorporating pre-collected offline data from a source environment can significantly improve the sample efficiency of reinforcement learning (RL), but this benefit is often challenged by discrepancies between the transition dynamics of the source and target environments. Existing methods typically a…

Cited by 0SourceScholar
2025

Context in Public Health for Underserved Communities: A Bayesian Approach to Online Restless Bandits

AAAI 2025technical

Public health programs often provide interventions to encourage program adherence, and effectively allocating interventions is vital for producing the greatest overall health outcomes, especially in underserved communities where resources are limited. Such resource allocation problems are often mode…

2025

Evaluating Index-based Treatment Allocation in Underresourced Communities

AAAI 2025technical

In many applications of AI for Social Impact (e.g., when allocating spots in support programs for underserved communities), resources are scarce and an allocation policy is needed to decide who receives a resource. Before being deployed at scale, a rigorous evaluation of an AI-powered allocation pol…

Cited by 0SourcePDFScholar
2025

Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement Learning

ICML 2025poster

In many real-world applications of Reinforcement Learning (RL), deployed policies have varied impacts on different stakeholders, creating challenges in reaching consensus on how to effectively aggregate their preferences. Generalized $p$-means form a widely used class of social welfare functions for…

Cited by 0SourcePDFScholar
2025

Optimizing Vital Sign Monitoring in Resource-Constrained Maternal Care: An RL-Based Restless Bandit Approach

AAAI 2025technical

Maternal mortality remains a significant global public health challenge. One promising approach to reducing maternal deaths occurring during facility-based childbirth is through early warning systems, which require the consistent monitoring of mothers' vital signs after giving birth. Wireless vital…

Cited by 3SourcePDFScholar
2025

PRIORITY2REWARD: Incorporating Healthworker Preferences for Resource Allocation Planning

AAAI 2025technical

In this paper, we present PRIORITY2REWARD a Large Language Model (LLM) based application which incorporates health worker preferences for resource allocation planning in public health programs. LLMs are increasingly used to design reward functions based on human preferences in Reinforcement Learning…

Cited by 0SourcePDFScholar
2025

Reinforcement learning with combinatorial actions for coupled restless bandits

ICLR 2025poster

Reinforcement learning (RL) has increasingly been applied to solve real-world planning problems, with progress in handling large state spaces and time horizons. However, a key bottleneck in many domains is that RL methods cannot accommodate large, combinatorially structured action spaces. In such se…

2025

Robust Optimization with Diffusion Models for Green Security

UAI 2025

In green security, defenders must forecast adversarial behavior-such as poaching, illegal logging, and illegal fishing-to plan effective patrols. These behavior are often highly uncertain and complex. Prior work has leveraged game theory to design robust patrol strategies to handle uncertainty, but

Cited by 0SourcePDFScholar
2025

The Bandit Whisperer: Communication Learning for Restless Bandits

AAAI 2025technical

Applying Reinforcement Learning (RL) to Restless Multi-Arm Bandits (RMABs) offers a promising avenue for addressing allocation problems with resource constraints and temporal dynamics. However, classic RMAB models largely overlook the challenges of (systematic) data errors - a common occurrence in r…

Cited by 6SourcePDFScholar
2025

What is the Right Notion of Distance between Predict-then-Optimize Tasks?

UAI 2025

Comparing datasets is a fundamental task in machine learning, essential for various learning paradigms-from evaluating train and test datasets for model generalization to using dataset similarity for detecting data drift. While traditional notions of dataset distances offer principled measures of si

2024

A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public Health

NeurIPS 2024poster

Restless multi-armed bandits (RMAB) have demonstrated success in optimizing resource allocation for large beneficiary populations in public health settings. Unfortunately, RMAB models lack flexibility to adapt to evolving public health policy priorities. Concurrently, Large Language Models (LLMs) ha…

Cited by 14SourcePDFScholar
2024

Group Fairness in Predict-Then-Optimize Settings for Restless Bandits

UAI 2024poster

Restless multi-arm bandits (RMABs) are a model for sequentially allocating a limited number of resources to agents modeled as Markov Decision Processes. RMABs have applications in cellular networks, anti-poaching, and in particular, healthcare. For such high-stakes use cases, allocations are often r…

Cited by 8SourcePDFScholar
2024

Improving Health Information Access in the World’s Largest Maternal Mobile Health Program via Bandit Algorithms

AAAI 2024technical

Harnessing the wide-spread availability of cell phones, many nonprofits have launched mobile health (mHealth) programs to deliver information via voice or text to beneficiaries in underserved communities, with maternal and infant health being a key area of such mHealth programs. Unfortunately, dwind…

Cited by 1SourcePDFScholar
2024

Leaving the Nest: Going beyond Local Loss Functions for Predict-Then-Optimize

AAAI 2024technical

Predict-then-Optimize is a framework for using machine learning to perform decision-making under uncertainty. The central research question it asks is, "How can we use the structure of a decision-making task to tailor ML models for that specific task?" To this end, recent work has proposed learning…

Cited by 14SourcePDFScholar
2024

Position: Application-Driven Innovation in Machine Learning

ICML 2024poster

In this position paper, we argue that application-driven research has been systemically under-valued in the machine learning community. As applications of machine learning proliferate, innovative algorithms inspired by specific real-world challenges have become increasingly important. Such work offe…

Cited by 4SourcePDFScholar
2024

Position: Social Environment Design Should be Further Developed for AI-based Policy-Making

ICML 2024poster

Artificial Intelligence (AI) holds promise as a technology that can be used to improve government and economic policy-making. This paper proposes a new research agenda towards this end by introducing **Social Environment Design**, a general framework for the use of AI in automated policy-making that…

Cited by 5SourcePDFScholar
2024

Towards a Pretrained Model for Restless Bandits via Multi-arm Generalization

IJCAI 2024poster

Restless multi-arm bandits (RMABs) is a class of resource allocation problems with broad application in areas such as healthcare, online advertising, and anti-poaching. We explore several important question such as how to handle arms opting-in and opting-out over time without frequent retraining fro…

2024

Transcendence: Generative Models Can Outperform The Experts That Train Them

NeurIPS 2024poster

Generative models are trained with the simple objective of imitating the conditional probability distribution induced by the data they are trained on. Therefore, when trained on data generated by humans, we may not expect the artificial model to outperform the humans on their original objectives. In…

Cited by 11SourcePDFScholar
2023

Complex Contagion Influence Maximization: A Reinforcement Learning Approach

IJCAI 2023poster

In influence maximization (IM), the goal is to find a set of seed nodes in a social network that maximizes the influence spread. While most IM problems focus on classical influence cascades (e.g., Independent Cascade and Linear Threshold) which assume individual influence cascade probability is inde…

Cited by 1SourcePDFScholar
2023

Find Rhinos without Finding Rhinos: Active Learning with Multimodal Imagery of South African Rhino Habitats

IJCAI 2023poster

Much of Earth's charismatic megafauna is endangered by human activities, particularly the rhino, which is at risk of extinction due to the poaching crisis in Africa. Monitoring rhinos' movement is crucial to their protection but has unfortunately proven difficult because rhinos are elusive. Therefor…

2023

Flexible Budgets in Restless Bandits: A Primal-Dual Algorithm for Efficient Budget Allocation

AAAI 2023technical

Restless multi-armed bandits (RMABs) are an important model to optimize allocation of limited resources in sequential decision-making settings. Typical RMABs assume the budget --- the number of arms pulled --- to be fixed for each step in the planning horizon. However, for realistic real-world plann…

Cited by 8SourcePDFScholar
2023

Improved Policy Evaluation for Randomized Trials of Algorithmic Resource Allocation

ICML 2023poster

We consider the task of evaluating policies of algorithmic resource allocation through randomized controlled trials (RCTs). Such policies are tasked with optimizing the utilization of limited intervention resources, with the goal of maximizing the benefits derived. Evaluation of such allocation poli…

Cited by 6SourcePDFScholar
2023

Increasing Impact of Mobile Health Programs: SAHELI for Maternal and Child Care

AAAI 2023technical

Underserved communities face critical health challenges due to lack of access to timely and reliable information. Nongovernmental organizations are leveraging the widespread use of cellphones to combat these healthcare challenges and spread preventative awareness. The health workers at these organiz…

2023

Limited Resource Allocation in a Non-Markovian World: The Case of Maternal and Child Healthcare

IJCAI 2023poster

The success of many healthcare programs depends on participants' adherence. We consider the problem of scheduling interventions in low resource settings (e.g., placing timely support calls from health workers) to increase adherence and/or engagement. Past works have successfully developed several cl…

Cited by 8SourcePDFScholar
2023

Optimistic Whittle Index Policy: Online Learning for Restless Bandits

AAAI 2023technical

Restless multi-armed bandits (RMABs) extend multi-armed bandits to allow for stateful arms, where the state of each arm evolves restlessly with different transitions depending on whether that arm is pulled. Solving RMABs requires information on transition dynamics, which are often unknown upfront. T…

2023

Robust Planning over Restless Groups: Engagement Interventions for a Large-Scale Maternal Telehealth Program

AAAI 2023technical

In 2020, maternal mortality in India was estimated to be as high as 130 deaths per 100K live births, nearly twice the UN's target. To improve health outcomes, the non-profit ARMMAN sends automated voice messages to expecting and new mothers across India. However, 38% of mothers stop listening to the…

Cited by 10SourcePDFScholar
2023

Scalable Decision-Focused Learning in Restless Multi-Armed Bandits with Application to Maternal and Child Health

AAAI 2023technical

This paper studies restless multi-armed bandit (RMAB) problems with unknown arm transition dynamics but with known correlated arm features. The goal is to learn a model to predict transition dynamics given features, where the Whittle index policy solves the RMAB problems using predicted transitions.…

Cited by 30SourcePDFScholar
2022

ADVISER: AI-Driven Vaccination Intervention Optimiser for Increasing Vaccine Uptake in Nigeria

IJCAI 2022poster

More than 5 million children under five years die from largely preventable or treatable medical conditions every year, with an overwhelmingly large proportion of deaths occurring in under-developed countries with low vaccination uptake. One of the United Nations' sustainable development goals (SDG 3…

2022

Coordinating Followers to Reach Better Equilibria: End-to-End Gradient Descent for Stackelberg Games

AAAI 2022technical

A growing body of work in game theory extends the traditional Stackelberg game to settings with one leader and multiple followers who play a Nash equilibrium. Standard approaches for computing equilibria in these games reformulate the followers' best response as constraints in the leader's optimizat…

Cited by 31SourcePDFScholar
2022

Decision-Focused Learning without Decision-Making: Learning Locally Optimized Decision Losses

NeurIPS 2022accept

Decision-Focused Learning (DFL) is a paradigm for tailoring a predictive model to a downstream optimization task that uses its predictions in order to perform better \textit{on that specific task}. The main technical challenge associated with DFL is that it requires being able to differentiate throu…

Cited by 52SourcePDFScholar
2022

Evolutionary Approach to Security Games with Signaling

IJCAI 2022poster

Green Security Games have become a popular way to model scenarios involving the protection of natural resources, such as wildlife. Sensors (e.g. drones equipped with cameras) have also begun to play a role in these scenarios by providing real-time information. Incorporating both human and sensor de…

2022

Ranked Prioritization of Groups in Combinatorial Bandit Allocation

IJCAI 2022poster

Preventing poaching through ranger patrols protects endangered wildlife, directly contributing to the UN Sustainable Development Goal 15 of life on land. Combinatorial bandits have been used to allocate limited patrol resources, but existing approaches overlook the fact that each location is home to…

2022

Restless and uncertain: Robust policies for restless bandits via deep multi-agent reinforcement learning

UAI 2022poster

We introduce robustness in \textit{restless multi-armed bandits} (RMABs), a popular model for constrained resource allocation among independent stochastic processes (arms). Nearly all RMAB techniques assume stochastic dynamics are precisely known. However, in many real-world settings, dynamics are e…

2022

Solving structured hierarchical games using differential backward induction

UAI 2022poster

From large-scale organizations to decentralized political systems, hierarchical strategic decision making is commonplace. We introduce a novel class of structured hierarchical games (SHGs) that formally capture such hierarchical strategic interactions. In an SHG, each player is a node in a tree, and…

2021

Contingency-aware influence maximization: A reinforcement learning approach

UAI 2021poster

The influence maximization (IM) problem aims at finding a subset of seed nodes in a social network that maximize the spread of influence. In this study, we focus on a sub-class of IM problems, where whether the nodes are willing to be the seeds when being invited is uncertain, called contingency-awa…

2021

Fair Influence Maximization: a Welfare Optimization Approach

AAAI 2021technical

Several behavioral, social, and public health interventions, such as suicide/HIV prevention or community preparedness against natural disasters, leverage social network information to maximize outreach. Algorithmic influence maximization techniques have been proposed to aid with the choice of ``peer…

Cited by 65SourcePDFScholar
2021

Learn to Intervene: An Adaptive Learning Policy for Restless Bandits in Application to Preventive Healthcare

IJCAI 2021poster

In many public health settings, it is important for patients to adhere to health programs, such as taking medications and periodic health checks. Unfortunately, beneficiaries may gradually disengage from such programs, which is detrimental to their health. A concrete example of gradual disengagement…

Cited by 58SourcePDFScholar
2021

Learning MDPs from Features: Predict-Then-Optimize for Sequential Decision Making by Reinforcement Learning

NeurIPS 2021spotlight

In the predict-then-optimize framework, the objective is to train a predictive model, mapping from environment features to parameters of an optimization problem, which maximizes decision quality when the optimization is subsequently solved. Recent work on decision-focused learning shows that embeddi…

Cited by 38SourcePDFScholar
2021

Robust reinforcement learning under minimax regret for green security

UAI 2021poster

Green security domains feature defenders who plan patrols in the face of uncertainty about the adversarial behavior of poachers, illegal loggers, and illegal fishers. Importantly, the deterrence effect of patrols on adversaries’ future behavior makes patrol planning a sequential decision-making prob…

2021

Tracking Disease Outbreaks from Sparse Data with Bayesian Inference

AAAI 2021technical

The COVID-19 pandemic provides new motivation for a classic problem in epidemiology: estimating the empirical rate of transmission during an outbreak (formally, the time-varying reproduction number) from case counts. While standard methods exist, they work best at coarse-grained national or state sc…

2020

Automatically Learning Compact Quality-aware Surrogates for Optimization Problems

NeurIPS 2020spotlight

Solving optimization problems with unknown parameters often requires learning a predictive model to predict the values of the unknown parameters and then solving the problem using these values. Recent work has shown that including the optimization problem as a layer in the model training pipeline re…

2020

Collapsing Bandits and Their Application to Public Health Intervention

NeurIPS 2020poster

We propose and study Collapsing Bandits, a new restless multi-armed bandit (RMAB) setting in which each arm follows a binary-state Markovian process with a special structure: when an arm is played, the state is fully observed, thus“collapsing” any uncertainty, but when an arm is passive, no observa…

2020

Robust Spatial-Temporal Incident Prediction

UAI 2020poster

Spatio-temporal incident prediction is a central issue in law enforcement, with applications in fighting crimes like poaching, human trafficking, illegal fishing, burglaries and smuggling. However, state of the art approaches fail to account for evasion in response to predictive models, a common fo…

Cited by 6SourcePDFScholar
2019

End to end learning and optimization on graphs

NeurIPS 2019poster

Real-world applications often combine learning and optimization problems on graphs. For instance, our objective may be to cluster the graph in order to detect meaningful communities (or solve other common graph optimization problems such as facility location, maxcut, and so on). However, graphs or r…

2019

Exploring Algorithmic Fairness in Robust Graph Covering Problems

NeurIPS 2019poster

Fueled by algorithmic advances, AI algorithms are increasingly being deployed in settings subject to unanticipated challenges with complex social effects. Motivated by real-world deployment of AI driven, social-network based suicide prevention and landslide risk management interventions, this paper…