← Search

Guojian Zhan

11 accepted papers

2026

DADP: Domain Adaptive Diffusion Policy

ICML 2026poster

Learning domain adaptive policies that can generalize to unseen transition dynamics, remains a fundamental challenge in learning-based control. Substantial progress has been made through domain representation learning to capture domain-specific information, thus enabling domain-aware decision making…

Cited by 0SourceScholar
2026

Harmonized Dual Policy Improvement for Model-based Reinforcement Learning

ICML 2026poster

Policy-planner bootstrapping has emerged as a powerful paradigm in model-based reinforcement learning (MBRL). We formalize this process as a dual policy improvement mechanism synergizing: (i) exploitative improvement via off-policy $Q$-maximization, and (ii) lookahead improvement via planner alignme…

Cited by 0SourceScholar
2026

Langevin Rollout Optimization for Modelic Reinforcement Learning

ICML 2026poster

Planning-driven model-based (modelic) reinforcement learning has achieved impressive success in continuous control tasks but predominantly relies on zero-order optimizers like Model Predictive Path Integral (MPPI). While robust for global exploration, MPPI updates actions solely through sampling and…

Cited by 0SourceScholar
2026

Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation

ICLR 2026oral

Learning expressive and efficient policy functions is a promising direction in reinforcement learning (RL). While flow-based policies have recently proven effective in modeling complex action distributions with a fast deterministic sampling process, they still face a trade-off between expressiveness…

Cited by 0SourceScholar
2026

Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL

ICML 2026poster

Maximum entropy has become a mainstream off-policy reinforcement learning (RL) framework for balancing exploitation and exploration. However, two bottlenecks still limit further performance gains: (1) non-stationary Q-value estimation stemming from the joint injection of entropy and the concurrent u…

Cited by 0SourceScholar
2026

Taming the Aleatoric Impulse in Off-Policy Reinforcement Learning

ICML 2026poster

Off-policy reinforcement learning is vulnerable to overestimation bias, which is rooted in the total value uncertainty. However, existing methods typically misaddress this by targeting the epistemic component, neglecting the aleatoric component. We identify for the first time that this oversight fai…

Cited by 0SourceScholar
2025

Bootstrap Off-policy with World Model

NeurIPS 2025poster

Online planning has proven effective in reinforcement learning (RL) for improving sample efficiency and final performance. However, using planning for environment interaction inevitably introduces a divergence between the collected data and the policy's actual behaviors, degrading both model learnin…

Cited by 0SourcecodeScholar
2025

Off-policy Reinforcement Learning with Model-based Exploration Augmentation

NeurIPS 2025poster

Exploration is crucial in Reinforcement Learning (RL) as it enables the agent to understand the environment for better decision-making. Existing exploration methods fall into two paradigms: active exploration, which injects stochasticity into the policy but struggles in high-dimensional environments…

Cited by 0SourceScholar
2025

Physics Informed Neural Pose Estimation for Real-Time Shape Reconstruction of Soft Continuum Robots

RA-L 2025

Soft continuum robots are increasingly valued for their remarkable flexibility, but accurate shape reconstruction remains challenging due to their infinite degrees of freedom and high nonlinearity. Existing approaches often rely on either simplified-curvature statics equations for physical derivatio

Cited by 4SourceScholar
2025

Transferable Latent-To-Latent Locomotion Policy for Efficient and Versatile Motion Control of Diverse Legged Robots

IROS 2025

Reinforcement learning (RL) has demonstrated remarkable capability in acquiring robot skills, but learning each new skill still requires substantial data collection for training. The pretrain-and-finetune paradigm offers a promising approach for efficiently adapting to new robot entities and tasks.

Cited by 2SourceScholar
2024

Rocket Landing Control with Random Annealing Jump Start Reinforcement Learning

IROS 2024

Rocket recycling is a crucial pursuit in aerospace technology, aimed at reducing costs and environmental impact in space exploration. The primary focus centers on rocket landing control, involving the guidance of a nonlinear under-actuated rocket with limited fuel in real-time. This challenging task

Cited by 6SourceScholar