← Search

Yanchi Ru

1 accepted papers

2026

RMO: Towards Better LLM Alignment via Reshaping Reward Margin Distributions

AAAI 2026technical

Large Language Models (LLMs) have achieved remarkable success in instruction-following and dialogue tasks, yet aligning them with human preferences remains a critical challenge. Recent advances such as Direct Preference Optimization (DPO) simplify the alignment pipeline by bypassing explicit reward

Cited by 0SourcePDFScholar