2024
Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning
NeurIPS 2024poster
This paper studies off-policy evaluation (OPE) in the presence of unmeasured confounders. Inspired by the two-way fixed effects regression model widely used in the panel data literature, we propose a two-way unmeasured confounding assumption to model the system dynamics in causal reinforcement learn…