← Search

Ruibo Deng

1 accepted papers

2026

AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment

AAAI 2026technical

Offline preference optimization offers a simpler and more stable alternative to RLHF for aligning language models. However, their effectiveness is critically dependent on ranking accuracy, a metric where further gains are highly impactful. This limitation arises from a fundamental problem that we id

Cited by 0SourcePDFScholar