← Search

Peihong Yu

3 accepted papers

2025

On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning

NeurIPS 2025poster

Reinforcement learning with general utilities (RLGU) offers a unifying framework to capture several problems beyond standard expected returns, including imitation learning, pure exploration, and safe RL. Despite recent fundamental advances in the theoretical analysis of policy gradient (PG) methods…

Cited by 0SourceScholar
2025

Sketch-to-Skill: Bootstrapping Robot Learning with Human Drawn Trajectory Sketches

RSS 2025poster

Training robotic manipulation policies traditionally requires numerous demonstrations and/or environmental rollouts. While recent Imitation Learning (IL) and Reinforcement Learning (RL) methods have reduced the number of required demonstrations, they still rely on expert knowledge to collect high-qu…

Cited by 1PDFScholar