← Search

Arpan Mukherjee

5 accepted papers

2026

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences

AAAI 2026technical

Uniform-reward reinforcement learning from human feedback (RLHF), which trains a single reward model to represent the preferences of all annotators, fails to capture the diversity of opinions across sub-populations, inadvertently favoring dominant groups. The state-of-the-art, MaxMin-RLHF, addresses

Cited by 0SourcePDFScholar
2026

Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality

ICLR 2026poster

While test-time scaling with verification has shown promise in improving the performance of large language models (LLMs), role of the verifier and its imperfections remain underexplored. The effect of verification manifests through interactions of three quantities: (i) the generator’s *coverage*, (i…

Cited by 0SourcecodeScholar
2025

Risk-sensitive Bandits: Arm Mixture Optimality and Regret-efficient Algorithms

AISTATS 2025poster

This paper introduces a general framework for risk-sensitive bandits that integrates the notions of risk-sensitive objectives by adopting a rich class of {\em distortion riskmetrics}. The introduced framework subsumes the various existing risk-sensitive models. An important and hitherto unknown obse…

Cited by 0SourcecodeScholar
2021

Mean-based Best Arm Identification in Stochastic Bandits under Reward Contamination

NeurIPS 2021poster

This paper investigates the problem of best arm identification in {\sl contaminated} stochastic multi-arm bandits. In this setting, the rewards obtained from any arm are replaced by samples from an adversarial model with probability $\varepsilon$. A fixed confidence (infinite-horizon) setting is con…

Cited by 14SourcePDFScholar