← Search

Daolong An

2 accepted papers

2026

State Proficiency-Based Adaptive Fine-Tuning for Offline-to-Online Reinforcement Learning

AAAI 2026technical

In offline-to-online (O2O) reinforcement learning, achieving efficient performance improvement while maintaining training stability remains a critical challenge for effective fine-tuning. Existing O2O methods usually focus on the balance between policy improvement and policy constraint during online

Cited by 0SourcePDFScholar
2024

Double Buffers CEM-TD3: More Efficient Evolution and Richer Exploration

AAAI 2024technical

CEM-TD3 is a combination scheme using the simple cross-entropy method (CEM) and Twin Delayed Deep Deterministic policy gradient (TD3), and it achieves a satisfactory trade-off between performance and sample efficiency. However, we find that CEM-TD3 cannot fully address the low efficiency of policy s…