← Search

Zhaoyang Gui

2 accepted papers

2025

Uncertainty-Aware Iterative Preference Optimization for Enhanced LLM Reasoning

ACL 2025long

Direct Preference Optimization (DPO) has recently emerged as an efficient and effective method for aligning large language models with human preferences. However, constructing high-quality preference datasets remains challenging, often necessitating expensive manual or powerful LM annotations. Addit…

2021

Adaptive Multi-Domain Learning for Outdoor 3d Human Pose and Shape Estimation

ICASSP 2021accepted

It is an extremely challenging task to estimate 3D human pose and shape in outdoor scenes for which we can hardly obtain precise ground truth data for training. Previous methods usually use multiple datasets collected at different scenes to train their models, including those collected in laboratori…

Cited by 0SourceScholar