← Search

Eric Mazumdar

21 accepted papers

2026

Distributionally Robust Cooperative Multi-agent Reinforcement Learning with Value Factorization

ICLR 2026poster

Cooperative multi-agent reinforcement learning (MARL) commonly adopts centralized training with decentralized execution, where value-factorization methods enforce the individual-global-maximum (IGM) principle so that decentralized greedy actions recover the team-optimal joint action. However, the re…

Cited by 0SourceScholar
2025

Breaking the Curse of Multiagency in Robust Multi-Agent Reinforcement Learning

ICML 2025poster

Standard multi-agent reinforcement learning (MARL) algorithms are vulnerable to sim-to-real gaps. To address this, distributionally robust Markov games (RMGs) have been proposed to enhance robustness in MARL by optimizing the worst-case performance when game dynamics shift within a prescribed uncert…

Cited by 5SourcePDFScholar
2025

Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning

ICLR 2025poster

Driven by inherent uncertainty and the sim-to-real gap, robust reinforcement learning (RL) seeks to improve resilience against the complexity and variability in agent-environment sequential interactions. Despite the existence of a large number of RL benchmarks, there is a lack of standardized benchm…

Cited by 1SourcePDFScholar
2025

Tractable Multi-Agent Reinforcement Learning through Behavioral Economics

ICLR 2025oral

A significant roadblock to the development of principled multi-agent reinforcement learning (MARL) algorithms is the fact that desired solution concepts like Nash equilibria may be intractable to compute. We show how one can overcome this obstacle by introducing concepts from behavioral economics in…

Cited by 0SourcePDFScholar
2024

Last-Iterate Convergence for Generalized Frank-Wolfe in Monotone Variational Inequalities

NeurIPS 2024poster

We study the convergence behavior of a generalized Frank-Wolfe algorithm in constrained (stochastic) monotone variational inequality (MVI) problems. In recent years, there have been numerous efforts to design algorithms for solving constrained MVI problems due to their connections with optimization,…

Cited by 0SourcePDFScholar
2024

Model-Free Robust $\phi$-Divergence Reinforcement Learning Using Both Offline and Online Data

ICML 2024poster

The robust $\phi$-regularized Markov Decision Process (RRMDP) framework focuses on designing control policies that are robust against parameter uncertainties due to mismatches between the simulator (nominal) model and real-world settings. This work makes *two* important contributions. First, we prop…

Cited by 5SourcePDFScholar
2024

Near-Optimal Distributionally Robust Reinforcement Learning with General $L_p$ Norms

NeurIPS 2024poster

To address the challenges of sim-to-real gap and sample efficiency in reinforcement learning (RL), this work studies distributionally robust Markov decision processes (RMDPs) --- optimize the worst-case performance when the deployed environment is within an uncertainty set around some nominal MDP. D…

Cited by 0SourcePDFScholar
2024

Sample-Efficient Robust Multi-Agent Reinforcement Learning in the Face of Environmental Uncertainty

ICML 2024poster

To overcome the sim-to-real gap in reinforcement learning (RL), learned policies must maintain robustness against environmental uncertainties. While robust RL has been widely studied in single-agent regimes, in multi-agent environments, the problem remains understudied---despite the fact that the pr…

Cited by 13SourcePDFScholar
2023

A Finite-Sample Analysis of Payoff-Based Independent Learning in Zero-Sum Stochastic Games

NeurIPS 2023poster

In this work, we study two-player zero-sum stochastic games and develop a variant of the smoothed best-response learning dynamics that combines independent learning dynamics for matrix games with the minimax value iteration for stochastic games. The resulting learning dynamics are payoff-based, conv…

Cited by 14SourcePDFScholar
2023

Algorithmic Collective Action in Machine Learning

ICML 2023poster

We initiate a principled study of algorithmic collective action on digital platforms that deploy machine learning algorithms. We propose a simple theoretical model of a collective interacting with a firm's learning algorithm. The collective pools the data of participating individuals and executes an…

Cited by 23SourcePDFScholar
2023

Strategic Distribution Shift of Interacting Agents via Coupled Gradient Flows

NeurIPS 2023poster

We propose a novel framework for analyzing the dynamics of distribution shift in real-world systems that captures the feedback loop between learning algorithms and the distributions on which they are deployed. Prior work largely models feedback-induced distribution shift as adversarial or via an ove…

Cited by 6SourcePDFScholar
2022

Decentralized, Communication- and Coordination-free Learning in Structured Matching Markets

NeurIPS 2022accept

We study the problem of online learning in competitive settings in the context of two-sided matching markets. In particular, one side of the market, the agents, must learn about their preferences over the other side, the firms, through repeated interaction while competing with other agents for succe…

Cited by 19SourcePDFScholar
2022

Zeroth-Order Methods for Convex-Concave Min-max Problems: Applications to Decision-Dependent Risk Minimization

AISTATS 2022poster

Min-max optimization is emerging as a key framework for analyzing problems of robustness to strategically and adversarially generated data. We propose the random reshuffling-based gradient-free Optimistic Gradient Descent-Ascent algorithm for solving convex-concave min-max problems with finite sum s…

Cited by 22SourcePDFScholar
2021

Global Convergence to Local Minmax Equilibrium in Classes of Nonconvex Zero-Sum Games

NeurIPS 2021poster

We study gradient descent-ascent learning dynamics with timescale separation ($\tau$-GDA) in unconstrained continuous action zero-sum games where the minimizing player faces a nonconvex optimization problem and the maximizing player optimizes a Polyak-Lojasiewicz (PL) or strongly-concave (SC) object…

Cited by 36SourcePDFScholar
2021

Who Leads and Who Follows in Strategic Classification?

NeurIPS 2021poster

As predictive models are deployed into the real world, they must increasingly contend with strategic behavior. A growing body of work on strategic classification treats this problem as a Stackelberg game: the decision-maker "leads" in the game by deploying a model, and the strategic agents "follow"…

Cited by 68SourcePDFScholar
2020

Feedback Linearization for Uncertain Systems via Reinforcement Learning

ICRA 2020poster

We present a novel approach to control design for nonlinear systems which leverages model-free policy optimization techniques to learn a linearizing controller for a physical plant with unknown dynamics. Feedback linearization is a technique from nonlinear control which renders the input-output dyna…

Cited by 49SourceScholar
2020

On Approximate Thompson Sampling with Langevin Algorithms

ICML 2020poster

Thompson sampling for multi-armed bandit problems is known to enjoy favorable performance in both theory and practice. However, its wider deployment is restricted due to a significant computational limitation: the need for samples from posterior distributions at every iteration. In practice, this li…

Cited by 40SourcePDFScholar
2019

Convergence Analysis of Gradient-Based Learning in Continuous Games

UAI 2019poster

Considering a class of gradient-based multi-agent learning algorithms in non-cooperative settings, we provide convergence guarantees to a neighborhood of a stable Nash equilibrium. In particular, we consider continuous games where agents learn in 1) deterministic settings with oracle access to their…

Cited by 40SourcePDFScholar