Online Pre-Training for Offline-to-Online Reinforcement Learning
Yongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngsoo Jang, Geon-Hyeong Kim, Jongseong Chae
Abstract
Offline-to-online reinforcement learning (RL) aims to integrate the complementary strengths of offline and online RL by pre-training an agent offline and subsequently fine-tuning it through online interactions. However, recent studies reveal that offline pre-trained agents often underperform during online fine-tuning due to inaccurate value estimation caused by distribution shift, with random initialization proving more effective in certain cases. In this work, we propose a novel method, Online Pre-Training for Offline-to-Online RL (OPT), explicitly designed to address the issue of inaccurate value estimation in offline pre-trained agents. OPT introduces a new learning phase, Online Pre-Training, which allows the training of a new value function tailored specifically for effective online fine-tuning. Implementation of OPT on TD3 and SPOT demonstrates an average 30\% improvement in performance across a wide range of D4RL environments, including MuJoCo, Antmaze, and Adroit.
BibTeX
@inproceedings{
shin2025online,
title={Online Pre-Training for Offline-to-Online Reinforcement Learning},
author={Yongjae Shin and Jeonghye Kim and Whiyoung Jung and Sunghoon Hong and Deunsol Yoon and Youngsoo Jang and Geon-Hyeong Kim and Jongseong Chae and Youngchul Sung and Kanghoon Lee and Woohyung Lim},
booktitle={Forty-second International Conference on Machine Learning},
year={2025},
url={https://openreview.net/forum?id=5anr1H5vR5}
}