← Search

Longxiang He

2 accepted papers

2025

FOSP: Fine-tuning Offline Safe Policy through World Models

ICLR 2025poster

Offline Safe Reinforcement Learning (RL) seeks to address safety constraints by learning from static datasets and restricting exploration. However, these approaches heavily rely on the dataset and struggle to generalize to unseen scenarios safely. In this paper, we aim to improve safety during the d…

2025

Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption

NeurIPS 2025poster

Pretraining a policy on offline data followed by fine-tuning through online interactions, known as Offline-to-Online Reinforcement Learning (O2O RL), has emerged as a promising paradigm for real-world RL deployment. However, both offline datasets and online interactions in practical environments are…

Cited by 0SourcecodeScholar