← Search

Hadi Khalaf

2 accepted papers

2026

Robust AI Evaluation through Maximal Lotteries

ICML 2026poster

The standard way to evaluate language models on subjective tasks is through pairwise comparisons: an annotator chooses the "better" of two model responses for a given prompt. These comparisons are then aggregated into a single ranking via the Bradley–Terry (BT) framework, forcing heterogeneous prefe…

Cited by 0SourceScholar
2025

Inference-Time Reward Hacking in Large Language Models

NeurIPS 2025spotlight

A common paradigm to improve the performance of large language models is optimizing for a reward model. Reward models assign a numerical score to an LLM’s output that indicates, for example, how likely it is to align with user preferences or safety goals. However, reward models are never perfect. Th…

Cited by 0SourceScholar