← Search

Debmalya Mandal

25 accepted papers

2026

BRIDGE: Bi-level Reinforcement Learning for Dynamic Group Structure in Coalition Formation Games

ICLR 2026poster

The challenge of coalition formation games lies in efficiently navigating the exponentially large space of possible coalitions to identify the optimal partition. While existing approaches to solve coalition formation games either provide optimal solutions with limited scalability or approximate solu…

Cited by 0SourceScholar
2025

Corruption Robust Offline Reinforcement Learning with Human Feedback

AISTATS 2025oral

We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedback about human preferences, an $\varepsilon$-fraction of the pairs is corrupted (e.g., feedback flipped or trajectory f…

Cited by 0SourceScholar
2025

Independent Learning in Performative Markov Potential Games

AISTATS 2025poster

Performative Reinforcement Learning (PRL) refers to a scenario in which the deployed policy changes the reward and transition dynamics of the underlying environment. In this work, we study multi-agent PRL by incorporating performative effects into Markov Potential Games (MPGs). We introduce the not…

Cited by 0SourcecodeScholar
2025

On Corruption-Robustness in Performative Reinforcement Learning

AAAI 2025technical

In performative Reinforcement Learning (RL), an agent faces a policy-dependent environment: the reward and transition functions depend on the agent's policy. Prior work on performative RL has studied the convergence of repeated retraining approaches to a performatively stable policy. In the finite s…

Cited by 1SourcePDFScholar
2025

Policy Teaching via Data Poisoning in Learning from Human Preferences

AISTATS 2025poster

We study data poisoning attacks in learning from human preferences. More specifically, we consider the problem of teaching/enforcing a target policy $\pi^\dagger$ by synthesizing preference data. We seek to understand the susceptibility of different preference-based learning paradigms to poisoned pr…

Cited by 0SourceScholar
2025

Stochastic Principal-Agent Problems: Computing and Learning Optimal History-Dependent Policies

NeurIPS 2025poster

We study a stochastic principal-agent model. A principal and an agent interact in a stochastic environment, each privy to observations about the state not available to the other. The principal has the power of commitment, both to elicit information from the agent and to signal her own information. T…

Cited by 0SourceScholar
2025

Strategyproof Reinforcement Learning from Human Feedback

NeurIPS 2025poster

We study Reinforcement Learning from Human Feedback (RLHF) in settings where multiple labelers may strategically misreport feedback to steer the learned policy toward their own preferences. We show that existing RLHF algorithms, including recent pluralistic methods, are not strategyproof, and that e…

Cited by 0SourceScholar
2024

Agent-Specific Effects: A Causal Effect Propagation Analysis in Multi-Agent MDPs

ICML 2024poster

Establishing causal relationships between actions and outcomes is fundamental for accountable multi-agent decision-making. However, interpreting and quantifying agents' contributions to such relationships pose significant challenges. These challenges are particularly prominent in the context of mult…

2024

Corruption-Robust Offline Two-Player Zero-Sum Markov Games

AISTATS 2024poster

We study data corruption robustness in offline two-player zero-sum Markov games. Given a dataset of realized trajectories of two players, an adversary is allowed to modify an $\epsilon$-fraction of it. The learner’s goal is to identify an approximate Nash Equilibrium policy pair from the corrupted d…

Cited by 3SourcePDFScholar
2024

Learning the Expected Core of Strictly Convex Stochastic Cooperative Games

NeurIPS 2024poster

Reward allocation, also known as the credit assignment problem, has been an important topic in economics, engineering, and machine learning. An important concept in reward allocation is the core, which is the set of stable allocations where no agent has the motivation to deviate from the grand coali…

2024

Performative Reinforcement Learning in Gradually Shifting Environments

UAI 2024poster

When Reinforcement Learning (RL) agents are deployed in practice, they might impact their environment and change its dynamics. We propose a new framework to model this phenomenon, where the current environment depends on the deployed policy as well as its previous dynamics. This is a generalization…

Cited by 6SourcePDFScholar
2024

Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences

ICML 2024poster

In this paper, we take a step towards a deeper understanding of learning from human preferences by systematically comparing the paradigm of reinforcement learning from human feedback (RLHF) with the recently proposed paradigm of direct preference optimization (DPO). We focus our attention on the cla…

Cited by 10SourcePDFScholar
2024

Symmetric Linear Bandits with Hidden Symmetry

NeurIPS 2024poster

High-dimensional linear bandits with low-dimensional structure have received considerable attention in recent studies due to their practical significance. The most common structure in the literature is sparsity. However, it may not be available in practice. Symmetry, where the reward is invariant un…

2023

Markov Decision Processes with Time-Varying Geometric Discounting

AAAI 2023technical

Canonical models of Markov decision processes (MDPs) usually consider geometric discounting based on a constant discount factor. While this standard modeling approach has led to many elegant results, some recent studies indicate the necessity of modeling time-varying discounting in certain applicati…

Cited by 2SourcePDFScholar
2023

Online Reinforcement Learning with Uncertain Episode Lengths

AAAI 2023technical

Existing episodic reinforcement algorithms assume that the length of an episode is fixed across time and known a priori. In this paper, we consider a general framework of episodic reinforcement learning when the length of each episode is drawn from a distribution. We first establish that this prob…

Cited by 7SourcePDFScholar
2021

Surprisingly Popular Voting Recovers Rankings, Surprisingly!

IJCAI 2021poster

The wisdom of the crowd has long become the de facto approach for eliciting information from individuals or experts in order to predict the ground truth. However, classical democratic approaches for aggregating individual \emph{votes} only work when the opinion of the majority of the crowd is relati…

Cited by 18SourcePDFScholar
2020

Ensuring Fairness Beyond the Training Data

NeurIPS 2020poster

We initiate the study of fair classifiers that are robust to perturbations in the training distribution. Despite recent progress, the literature on fairness has largely ignored the design of fair and robust classifiers. In this work, we develop classifiers that are fair not only with respect to the…

2019

Efficient and Thrifty Voting by Any Means Necessary

NeurIPS 2019oral

We take an unorthodox view of voting by expanding the design space to include both the elicitation rule, whereby voters map their (cardinal) preferences to votes, and the aggregation rule, which transforms the reported votes into collective decisions. Intuitively, there is a tradeoff between the com…

Cited by 62SourcePDFScholar