← Search

Sayak Ray Chowdhury

20 accepted papers

2025

Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift

ICML 2025poster

Current Large Language Model (LLM) preference optimization algorithms do not account for temporal preference drift, which can lead to severe misalignment. To address this limitation, we propose **Non-Stationary Direct Preference Optimisation (NS-DPO)** that models time-dependent reward functions wit…

Cited by 0SourcePDFScholar
2024

Differentially Private Reward Estimation with Preference Feedback

AISTATS 2024poster

Learning from preference-based feedback has recently gained considerable traction as a promising approach to align generative models with human interests. Instead of relying on numerical rewards, the generative models are trained using reinforcement learning with human feedback (RLHF). These approac…

Cited by 6SourcePDFScholar
2024

OAK: Enriching Document Representations using Auxiliary Knowledge for Extreme Classification

ICML 2024poster

The objective in eXtreme Classification (XC) is to find relevant labels for a document from an exceptionally large label space. Most XC application scenarios have rich auxiliary data associated with the input documents, e.g., frequently clicked webpages for search queries in sponsored search. Unfort…

Cited by 2SourcePDFScholar
2024

Provably Robust DPO: Aligning Language Models with Noisy Feedback

ICML 2024poster

Learning from preference-based feedback has recently gained traction as a promising approach to align language models with human interests. While these aligned generative models have demonstrated impressive capabilities across various tasks, their dependence on high-quality human preference data pos…

Cited by 53SourcePDFScholar
2023

Combinatorial categorized bandits with expert rankings

UAI 2023poster

Many real-world systems such as e-commerce websites and content-serving platforms employ two-stage recommendation — in the first stage, multiple nominators (experts) provide ranked lists of items (one nominator per category, e.g., sports and political news articles), and in the second stage, an aggr…

Cited by 2SourcePDFScholar
2023

Differentially Private Episodic Reinforcement Learning with Heavy-tailed Rewards

ICML 2023poster

In this paper we study the problem of (finite horizon tabular) Markov decision processes (MDPs) with heavy-tailed rewards under the constraint of differential privacy (DP). Compared with the previous studies for private reinforcement learning that typically assume rewards are sampled from some bound…

Cited by 1SourcePDFScholar
2023

Exploration in Linear Bandits with Rich Action Sets and its Implications for Inference

AISTATS 2023poster

We present a non-asymptotic lower bound on the spectrum of the design matrix generated by any linear bandit algorithm with sub-linear regret when the action set has well-behaved curvature. Specifically, we show that the minimum eigenvalue of the expected design matrix grows as $\Omega(\sqrt{n})$ whe…

Cited by 6SourcePDFScholar
2022

Differentially Private Regret Minimization in Episodic Markov Decision Processes

AAAI 2022technical

We study regret minimization in finite horizon tabular Markov decision processes (MDPs) under the constraints of differential privacy (DP). This is motivated by the widespread applications of reinforcement learning (RL) in real-world sequential decision making problems, where protecting users' sensi…

2021

Reinforcement Learning in Parametric MDPs with Exponential Families

AISTATS 2021poster

Extending model-based regret minimization strategies for Markov decision processes (MDPs) beyond discrete state-action spaces requires structural assumptions on the reward and transition models. Existing parametric approaches establish regret guarantees by making strong assumptions about either the…

Cited by 17SourcePDFScholar
2020

Active Learning of Conditional Mean Embeddings via Bayesian Optimisation

UAI 2020poster

We consider the problem of sequentially optimising the conditional expectation of an objective function, with both the conditional distribution and the objective function assumed to be fixed but unknown. Assuming that the objective function belongs to a reproducing kernel Hilbert space (RKHS), we pr…

Cited by 11SourcePDFScholar