2026
Reward-Preserving Counterfactual State Editing for Offline Reinforcement Learning
ICML 2026poster
Transformer sequence models such as Decision Transformer can learn strong offline policies from logged trajectories, but they can suffer from causal confusion: reliance on spurious correlations that predict reward in the data but do not reflect the true causal mechanisms of the environment. We propo…