← Search

Yucheng Yang

6 accepted papers

2026

GuidedBridge: Training-freely Improving Bridge Models with Prior Guidance

ICML 2026poster

Guidance methods, e.g., classifier-free guidance (CFG) and auto-guidance (AG), have distinctively improved noise-to-data diffusion generation results. Recently, bridge models have been proposed, which present a data-to-data sampling process to exploit instructive information from clean prior represe…

Cited by 0SourceScholar
2026

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

ICML 2026spotlight

Mean Field Games (MFGs) provide a principled framework for modeling interactions in large populations models: at scale, population dynamics become deterministic, with uncertainty entering only through aggregate shocks, or *common noise*. However, algorithmic progress has been limited since model-fre…

Cited by 0SourceScholar
2025

Preference Controllable Reinforcement Learning with Advanced Multi-Objective Optimization

ICML 2025poster

Practical reinforcement learning (RL) usually requires agents to be optimized for multiple potentially conflicting criteria, e.g. speed vs. safety. Although Multi-Objective RL (MORL) algorithms have been studied in previous works, their trained agents often cover limited Pareto optimal solutions an…

Cited by 0SourcePDFScholar
2024

Task Adaptation from Skills: Information Geometry, Disentanglement, and New Objectives for Unsupervised Reinforcement Learning

ICLR 2024spotlight

Unsupervised reinforcement learning (URL) aims to learn general skills for unseen downstream tasks. Mutual Information Skill Learning (MISL) addresses URL by maximizing the mutual information between states and skills but lacks sufficient theoretical analysis, e.g., how well its learned skills can i…

Cited by 7SourcePDFScholar
2022

Keypoint-Guided Optimal Transport with Applications in Heterogeneous Domain Adaptation

NeurIPS 2022accept

Existing Optimal Transport (OT) methods mainly derive the optimal transport plan/matching under the criterion of transport cost/distance minimization, which may cause incorrect matching in some cases. In many applications, annotating a few matched keypoints across domains is reasonable or even effor…

Cited by 34SourcePDFScholar