← Search

Kousha Kalantari

4 accepted papers

2025

Comparing Few to Rank Many: Active Human Preference Learning Using Randomized Frank-Wolfe Method

ICML 2025poster

We study learning human preferences from limited comparison feedback, a core machine learning problem that is at the center of reinforcement learning from human feedback (RLHF). We formulate the problem as learning a Plackett-Luce (PL) model from a limited number of $K$-subset comparisons over a uni…

Cited by 0SourcePDFScholar
2025

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain

ICML 2025poster

Supervised fine-tuning (SFT) is the most common way of adapting large language models (LLMs) to a new domain. In this paper, we improve the efficiency of SFT by selecting an informative subset of training examples. Specifically, for a fixed budget of training examples, which determines the computati…

Cited by 0SourcePDFScholar
2024

Optimal Design for Human Preference Elicitation

NeurIPS 2024poster

Learning of preference models from human feedback has been central to recent advances in artificial intelligence. Motivated by the cost of obtaining high-quality human annotations, we study efficient human preference elicitation for learning preference models. The key idea in our work is to generali…

Cited by 5SourcePDFScholar
2023

Fixed-Budget Best-Arm Identification with Heterogeneous Reward Variances

UAI 2023poster

We study the problem of best-arm identification (BAI) in the fixed-budget setting with heterogeneous reward variances. We propose two variance-adaptive BAI algorithms for this setting: SHVar for known reward variances and SHAdaVar for unknown reward variances. Our algorithms rely on non-uniform budg…

Cited by 9SourcePDFScholar