2026
Long-Horizon Model-Based Offline Reinforcement Learning Without Conservatism
ICML 2026poster
Popular offline reinforcement learning (RL) methods rely on conservatism, penalizing out-of-dataset actions or restricting rollout horizons. We question the universality of this principle and revisit a complementary Bayesian perspective. By modeling a posterior over plausible world models and traini…