← Search

Yubing Wu

1 accepted papers

2026

Towards Disentangled Preference Optimization Dynamics

ICML 2026poster

Preference optimization is widely used to align large language models (LLMs) with human preferences, yet many margin-based objectives often suppress the chosen response together with the rejected one, and no general mechanism exists to prevent this across objectives. We bridge this gap by presenting…

Cited by 0SourceScholar