← Search

Ruibo Guo

2 accepted papers

2026

Tackling Heavy-Tailed Q-Value Bias in Offline-to-Online Reinforcement Learning with Laplace-Robust Modeling

ICLR 2026poster

Offline-to-online reinforcement learning (O2O RL) aims to improve the performance of offline pretrained agents through online fine-tuning. Existing O2O RL methods have achieved advances in mitigating the overestimation of Q-value biases (i.e., biases of cumulative rewards), improving the performance…

Cited by 0SourceScholar
2025

Learning Robust Representations with Long-Term Information for Generalization in Visual Reinforcement Learning

ICLR 2025poster

Generalization in visual reinforcement learning (VRL) aims to learn agents that can adapt to test environments with unseen visual distractions. Despite advances in robust representations learning, many methods do not take into account the essential downstream task of sequential decision-making. This…

Cited by 0SourcePDFScholar