← Search

Shiyue Xu

1 accepted papers

2025

MWPO: Enhancing LLMs Performance through Multi-Weight Preference Strength and Length Optimization

ACL 2025finding

Direct Preference Optimization (DPO) have proposed offline alternatives to Reinforcement Learning from Human Feedback (RLHF). In DPO, each preference pair, which serves as the foundation for learning, is typically constructed by first generating multiple responses to the same instruction and then an…