← Search

Ali Aouad

1 accepted papers

2026

The Sign Estimator: Preference Modeling for LLM Alignment under Heterogeneity

ICML 2026poster

Traditional large language model (LLM) alignment methods are based on Reinforcement Learning From Human Feedback (RLHF), which learns a single reward model (implicitly or explicitly) from pairwise comparison data. This approach implicitly assumes homogeneous preferences across human labelers---an as…

Cited by 0SourceScholar