← Search

Sonali Parbhoo

10 accepted papers

2026

Reward Shaping Control Variates for Off-Policy Evaluation Under Sparse Rewards

ICML 2026poster

Off-policy evaluation (OPE) is essential for deploying reinforcement learning in safety-critical settings, yet existing estimators such as importance sampling and doubly robust (DR) often exhibit prohibitively high variance when rewards are sparse. In this work, we introduce Reward-Shaping Control V…

Cited by 0SourceScholar
2026

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

ICLR 2026poster

The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a grand challenge. While Inverse Reinforcement Learning (IRL) can infer reward functions from behaviour, existing approaches either produce a single, overconfiden…

Cited by 0SourceScholar
2025

Do Regularization Methods for Shortcut Mitigation Work As Intended?

AISTATS 2025poster

Mitigating shortcuts, where models exploit spurious correlations in training data, remains a significant challenge for improving generalization. Regularization methods have been proposed to address this issue by enhancing model generalizability. However, we demonstrate that these methods can sometim…

Cited by 0SourcecodeScholar
2023

The Unintended Consequences of Discount Regularization: Improving Regularization in Certainty Equivalence Reinforcement Learning

ICML 2023poster

Discount regularization, using a shorter planning horizon when calculating the optimal policy, is a popular choice to restrict planning to a less complex set of policies when estimating an MDP from sparse or noisy data (Jiang et al., 2015). It is commonly understood that discount regularization func…

Cited by 5SourcePDFScholar
2020

Interpretable Off-Policy Evaluation in Reinforcement Learning by Highlighting Influential Transitions

ICML 2020poster

Off-policy evaluation in reinforcement learning offers the chance of using observational data to improve future outcomes in domains such as healthcare and education, but safe deployment in high stakes settings requires ways of assessing its validity. Traditional measures such as confidence intervals…

2019

Greedy Structure Learning of Hierarchical Compositional Models

CVPR 2019poster

In this work, we consider the problem of learning a hierarchical generative model of an object from a set of images which show examples of the object in the presence of variable background clutter. Existing approaches to this problem are limited by making strong a-priori assumptions about the object…

Cited by 11PDFScholar
2016

Bayesian Markov Blanket Estimation

AISTATS 2016poster

This paper considers a Bayesian view for estimating the Markov blanket of a set of query variables, where the set of potential neighbours here is big. We factorize the posterior such that the Markov blanket is conditionally independent of the network of the potential neighbours. By exploiting this…

Cited by 8SourcePDFScholar