ICLR 2026poster0 citations

Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data

Yi Zhao, Aidan Scannell, Wenshuai Zhao, Yuxin Hou, Tianyu Cui, Le Chen, Dieter Büchler, Arno Solin

Abstract

Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL). This paper expands the pool of usable data for offline-to-online RL by leveraging abundant non-curated data that is reward-free, of mixed quality, and collected across multiple embodiments. Although learning a world model appears promising for utilizing such data, we find that naive fine-tuning fails to accelerate RL training on many tasks. Through careful investigation, we attribute this failure to the distributional shift between offline and online data during fine-tuning. To address this issue and effectively use the offline data, we propose two techniques: i) experience rehearsal and ii) execution guidance. With these modifications, the non-curated offline data substantially improves RL’s sample efficiency. Under limited sample budgets, our method achieves a 102.8% relative improvement in aggregate score over learning-from-scratch baselines across 72 visuomotor tasks spanning 6 embodiments. On challenging tasks such as locomotion and robotic manipulation, it outperforms prior methods that utilize offline data by a decent margin.

Reinforcement LearningReinforcement Learning from offline data
BibTeX
@inproceedings{
zhao2026efficient,
title={Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data},
author={Yi Zhao and Aidan Scannell and Wenshuai Zhao and Yuxin Hou and Tianyu Cui and Le Chen and Dieter B{\"u}chler and Arno Solin and Juho Kannala and Joni Pajarinen},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=oBXfPyi47m}
}
Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data · ICLR 2026