← Search

Yang Guan

5 accepted papers

2026

Harmonized Dual Policy Improvement for Model-based Reinforcement Learning

ICML 2026poster

Policy-planner bootstrapping has emerged as a powerful paradigm in model-based reinforcement learning (MBRL). We formalize this process as a dual policy improvement mechanism synergizing: (i) exploitative improvement via off-policy $Q$-maximization, and (ii) lookahead improvement via planner alignme…

Cited by 0SourceScholar
2026

Langevin Rollout Optimization for Modelic Reinforcement Learning

ICML 2026poster

Planning-driven model-based (modelic) reinforcement learning has achieved impressive success in continuous control tasks but predominantly relies on zero-order optimizers like Model Predictive Path Integral (MPPI). While robust for global exploration, MPPI updates actions solely through sampling and…

Cited by 0SourceScholar
2026

Taming the Aleatoric Impulse in Off-Policy Reinforcement Learning

ICML 2026poster

Off-policy reinforcement learning is vulnerable to overestimation bias, which is rooted in the total value uncertainty. However, existing methods typically misaddress this by targeting the epistemic component, neglecting the aleatoric component. We identify for the first time that this oversight fai…

Cited by 0SourceScholar
2024

VeraCT Scan: Retrieval-Augmented Fake News Detection with Justifiable Reasoning

ACL 2024system demonstrations

The proliferation of fake news poses a significant threat not only by disseminating misleading information but also by undermining the very foundations of democracy. The recent advance of generative artificial intelligence has further exacerbated the challenge of distinguishing genuine news from fab…

Cited by 2SourcePDFScholar
2021

Model-based Constrained Reinforcement Learning using Generalized Control Barrier Function

IROS 2021poster

Model information can be used to predict future trajectories, so it has huge potential to avoid dangerous regions when applying reinforcement learning (RL) on real-world tasks, like autonomous driving. However, existing studies mostly use model-free constrained RL, which causes inevitable constraint…

Cited by 86SourcecodeScholar