← Search

David Yao

4 accepted papers

2025

MallowsPO: Fine-Tune Your LLM with Preference Dispersions

ICLR 2025poster

Direct Preference Optimization (DPO) has recently emerged as a popular approach to improve reinforcement learning from human feedback (RLHF), leading to better techniques to fine-tune large language models (LLM). A weakness of DPO, however, lies in its lack of capability to characterize the diversit…

Cited by 6SourcePDFScholar
2025

RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization

ICLR 2025poster

Recently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully aligned models with human preferences, there is a lack of understanding regarding the contributions of their additional c…

Cited by 5SourcePDFScholar
2025

Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning

ICML 2025poster

Reinforcement learning from human feedback (RLHF), which aligns a diffusion model with input prompt, has become a crucial step in building reliable generative AI models. Most works in this area uses a *discrete-time* formulation, which is prone to induced errors, and often not applicable to models w…

Cited by 1SourcePDFScholar