← Search

Shengbo Li

9 accepted papers

2026

Harmonized Dual Policy Improvement for Model-based Reinforcement Learning

ICML 2026poster

Policy-planner bootstrapping has emerged as a powerful paradigm in model-based reinforcement learning (MBRL). We formalize this process as a dual policy improvement mechanism synergizing: (i) exploitative improvement via off-policy $Q$-maximization, and (ii) lookahead improvement via planner alignme…

Cited by 0SourceScholar
2026

Langevin Rollout Optimization for Modelic Reinforcement Learning

ICML 2026poster

Planning-driven model-based (modelic) reinforcement learning has achieved impressive success in continuous control tasks but predominantly relies on zero-order optimizers like Model Predictive Path Integral (MPPI). While robust for global exploration, MPPI updates actions solely through sampling and…

Cited by 0SourceScholar
2026

Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL

ICML 2026poster

Maximum entropy has become a mainstream off-policy reinforcement learning (RL) framework for balancing exploitation and exploration. However, two bottlenecks still limit further performance gains: (1) non-stationary Q-value estimation stemming from the joint injection of entropy and the concurrent u…

Cited by 0SourceScholar
2026

SwarmNav: Swarm Robotics Navigation in Dynamic and Dense Environments Via Reinforcement Learning

ICRA 2026poster

Collision avoidance and navigation in dynamic and dense environments remain highly challenging for swarm robotics. To address this, we propose SwarmNav, a novel goal-region amplification navigation policy that leverages LiDAR-based position data to generate velocity commands guiding robots toward th…

Cited by 0Scholar
2026

Taming the Aleatoric Impulse in Off-Policy Reinforcement Learning

ICML 2026poster

Off-policy reinforcement learning is vulnerable to overestimation bias, which is rooted in the total value uncertainty. However, existing methods typically misaddress this by targeting the epistemic component, neglecting the aleatoric component. We identify for the first time that this oversight fai…

Cited by 0SourceScholar
2025

FLARE: Fast Large-Scale Autonomous Exploration Guided by Unknown Regions

RA-L 2025

Autonomous exploration is a critical foundation for unmanned aerial vehicle (UAV) applications such as search and rescue. However, existing methods typically focus only on known spaces or frontiers without considering unknown regions or providing further guidance for the global path, which results i

Cited by 2SourceScholar
2022

Flow-based Recurrent Belief State Learning for POMDPs

ICML 2022spotlight

Partially Observable Markov Decision Process (POMDP) provides a principled and generic framework to model real world sequential decision making processes but yet remains unsolved, especially for high dimensional continuous space and unknown models. The main challenge lies in how to accurately obtain…

Cited by 25SourcePDFScholar