← Search

Mengyang Li

6 accepted papers

2026

Aligner, Diagnose Thyself: A Meta-Learning Paradigm for Fusing Intrinsic Feedback in Preference Alignment

ICLR 2026poster

The alignment of Large Language Models (LLMs) with human preferences is critically undermined by noisy labels in training datasets. Existing robust methods often prove insufficient, as they rely on single, narrow heuristics such as perplexity or loss, failing to address the diverse nature of real-wo…

Cited by 0SourceScholar
2026

Layer-wise Gradient Disentanglement: Decoupling Semantics and Preferences in Direct Preference Optimization

ICML 2026poster

Direct Preference Optimization (DPO) has become the dominant approach for aligning large language models with human preferences. However, standard DPO treats all preference pairs uniformly, overlooking the heterogeneous nature of the learning problem: some samples demand sophisticated semantic under…

Cited by 0SourceScholar