← Search

Susan Murphy

8 accepted papers

2025

Harnessing Causality in Reinforcement Learning with Bagged Decision Times

AISTATS 2025poster

We consider reinforcement learning (RL) for a class of problems with bagged decision times. A bag contains a finite sequence of consecutive decision times. The transition dynamics are non-Markovian and non-stationary within a bag. All actions within a bag jointly impact a single reward, observed at…

Cited by 0SourceScholar
2024

Contextual Bandits with Budgeted Information Reveal

AISTATS 2024poster

Contextual bandit algorithms are commonly used in digital health to recommend personalized treatments. However, to ensure the effectiveness of the treatments, patients are often requested to take actions that have no immediate benefit to them, which we refer to as pro-treatment actions. In practice,…

Cited by 5SourcePDFScholar
2024

ReBandit: Random Effects Based Online RL Algorithm for Reducing Cannabis Use

IJCAI 2024poster

The escalating prevalence of cannabis use, and associated cannabis-use disorder (CUD), poses a significant public health challenge globally. With a notably wide treatment gap, especially among emerging adults (EAs; ages 18-25), addressing cannabis use and CUD remains a pivotal objective within the 2…

2023

The Unintended Consequences of Discount Regularization: Improving Regularization in Certainty Equivalence Reinforcement Learning

ICML 2023poster

Discount regularization, using a shorter planning horizon when calculating the optimal policy, is a popular choice to restrict planning to a less complex set of policies when estimating an MDP from sparse or noisy data (Jiang et al., 2015). It is commonly understood that discount regularization func…

Cited by 5SourcePDFScholar
2021

Statistical Inference with M-Estimators on Adaptively Collected Data

NeurIPS 2021poster

Bandit algorithms are increasingly used in real-world sequential decision-making problems. Associated with this is an increased desire to be able to use the resulting datasets to answer scientific questions like: Did one type of ad lead to more purchases? In which contexts is a mobile health interve…

Cited by 64SourcePDFScholar