← Search

Ayush Sawarni

5 accepted papers

2025

Preference Learning with Response Time: Robust Losses and Guarantees

NeurIPS 2025poster

This paper investigates the integration of response time data into human preference learning frameworks for more effective reward model elicitation. While binary preference data has become fundamental in fine-tuning foundation models, generative AI systems, and other large-scale models, the valuable…

Cited by 0SourceScholar
2024

Generalized Linear Bandits with Limited Adaptivity

NeurIPS 2024spotlight

We study the generalized linear contextual bandit problem within the constraints of limited adaptivity. In this paper, we present two algorithms, B-GLinCB and RS-GLinCB, that address, respectively, two prevalent limited adaptivity settings. Given a budget $M$ on the number of policy updates, in the…

2023

Fairness and Welfare Quantification for Regret in Multi-Armed Bandits

AAAI 2023technical

We extend the notion of regret with a welfarist perspective. Focussing on the classic multi-armed bandit (MAB) framework, the current work quantifies the performance of bandit algorithms by applying a fundamental welfare function, namely the Nash social welfare (NSW) function. This corresponds to eq…

Cited by 18SourcePDFScholar
2023

Learning good interventions in causal graphs via covering

UAI 2023poster

We study the causal bandit problem that entails identifying a near-optimal intervention from a specified set A of (possibly non-atomic) interventions over a given causal graph. Here, an optimal intervention in A is one that maximizes the expected value for a designated reward variable in the graph,…