← Search

Hanshan Zhang

1 accepted papers

2025

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model

ICML 2025poster

Since the debut of DPO, it has been shown that aligning a target LLM with human preferences via the KL-constrained RLHF loss is mathematically equivalent to a special kind of reward modeling task. Concretely, the task requires: 1) using the target LLM to parameterize the reward model, and 2) tuning…

Cited by 1SourcePDFScholar