← Search

Shrey Modi

2 accepted papers

2025

Selective Uncertainty Propagation in Offline RL

AAAI 2025technical

We consider the finite-horizon offline reinforcement learning (RL) setting, and are motivated by the challenge of learning the policy at any step h in dynamic programming (DP) algorithms. To learn this, it is sufficient to evaluate the treatment effect of deviating from the behavioral policy at step…

Cited by 1SourcePDFScholar