← Search

Zhao Shan

3 accepted papers

2025

Forward KL Regularized Preference Optimization for Aligning Diffusion Policies

AAAI 2025technical

Diffusion models have achieved remarkable success in sequential decision-making by leveraging the highly expressive model capabilities in policy learning. A central problem for learning diffusion policies is to align the policy output with human intents in various tasks. To achieve this, previous me…

Cited by 3SourcePDFScholar
2025

Preference Aligned Diffusion Planner for Quadrupedal Locomotion Control

IROS 2025

Diffusion models demonstrate superior performance in capturing complex distributions from large-scale datasets, providing a promising solution for quadrupedal locomotion control. However, the robustness of the diffusion planner is inherently dependent on the diversity of the pre-collected datasets.

Cited by 9SourcecodeScholar
2025

Task-Agnostic Pre-training and Task-Guided Fine-tuning for Versatile Diffusion Planner

ICML 2025poster

Diffusion models have demonstrated their capabilities in modeling trajectories of multi-tasks. However, existing multi-task planners or policies typically rely on task-specific demonstrations via multi-task imitation, or require task-specific reward labels to facilitate policy optimization via Reinf…

Cited by 5SourcePDFScholar