NeurIPS 2024poster5 citations

Optimal Design for Human Preference Elicitation

Subhojyoti Mukherjee, Anusha Lalitha, Kousha Kalantari, Aniket Anand Deshmukh, Ge Liu, Yifei Ma, Branislav Kveton

Abstract

Learning of preference models from human feedback has been central to recent advances in artificial intelligence. Motivated by the cost of obtaining high-quality human annotations, we study efficient human preference elicitation for learning preference models. The key idea in our work is to generalize optimal designs, an approach to computing optimal information-gathering policies, to lists of items that represent potential questions with answers. The policy is a distribution over the lists and we elicit preferences from them proportionally to their probabilities. To show the generality of our ideas, we study both absolute and ranking feedback models on items in the list. We design efficient algorithms for both and analyze them. Finally, we demonstrate that our algorithms are practical by evaluating them on existing question-answering problems.

optimal designactive learningpreference elicitationhuman feedbacklearning to rank
BibTeX
@inproceedings{
mukherjee2024optimal,
title={Optimal Design for Human Preference Elicitation},
author={Subhojyoti Mukherjee and Anusha Lalitha and Kousha Kalantari and Aniket Anand Deshmukh and Ge Liu and Yifei Ma and Branislav Kveton},
booktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},
year={2024},
url={https://openreview.net/forum?id=cCGWj61Ael}
}