← Search

Shikai Luo

3 accepted papers

2024

Robust Offline Reinforcement Learning with Heavy-Tailed Rewards

AISTATS 2024poster

This paper endeavors to augment the robustness of offline reinforcement learning (RL) in scenarios laden with heavy-tailed rewards, a prevalent circumstance in real-world applications. We propose two algorithmic frameworks, ROAM and ROOM, for robust off-policy evaluation and offline policy optimizat…

2023

An Instrumental Variable Approach to Confounded Off-Policy Evaluation

ICML 2023poster

Off-policy evaluation (OPE) aims to estimate the return of a target policy using some pre-collected observational data generated by a potentially different behavior policy. In many cases, there exist unmeasured variables that confound the action-reward or action-next-state relationships, rendering m…

Cited by 21SourcePDFScholar