2026
RMO: Towards Better LLM Alignment via Reshaping Reward Margin Distributions
AAAI 2026technical
Large Language Models (LLMs) have achieved remarkable success in instruction-following and dialogue tasks, yet aligning them with human preferences remains a critical challenge. Recent advances such as Direct Preference Optimization (DPO) simplify the alignment pipeline by bypassing explicit reward