2026
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
ICLR 2026poster
The evaluation of large language models faces significant challenges. Technical benchmarks often lack real-world relevance, while existing human preference evaluations suffer from unrepresentative sampling, superficial assessment depth, and single-metric reductionism. To address these issues, we int…