← Search

Maryam Kamgarpour

19 accepted papers

2026

Nash Equilibria in Games with Playerwise Concave Coupling Constraints: Existence and Computation

ICML 2026oral

We study the existence and computation of Nash equilibria in concave games where the players' admissible strategies are subject to shared coupling constraints. Under playerwise concavity of constraints, we prove existence of Nash equilibria. Our proof leverages topological fixed point theory and nov…

Cited by 0SourceScholar
2025

Efficient Preference-Based Reinforcement Learning: Randomized Exploration meets Experimental Design

NeurIPS 2025poster

We study reinforcement learning from human feedback in general Markov decision processes, where agents learn from trajectory-level preference comparisons. A central challenge in this setting is to design algorithms that select informative preference queries to identify the underlying reward while en…

Cited by 0SourceScholar
2024

Towards the Transferability of Rewards Recovered via Regularized Inverse Reinforcement Learning

NeurIPS 2024poster

Inverse reinforcement learning (IRL) aims to infer a reward from expert demonstrations, motivated by the idea that the reward, rather than the policy, is the most succinct and transferable description of a task [Ng et al., 2000]. However, the reward corresponding to an optimal policy is not unique,…

Cited by 3SourcePDFScholar
2023

Identifiability and Generalizability in Constrained Inverse Reinforcement Learning

ICML 2023poster

Two main challenges in Reinforcement Learning (RL) are designing appropriate reward functions and ensuring the safety of the learned policy. To address these challenges, we present a theoretical framework for Inverse Reinforcement Learning (IRL) in constrained Markov decision processes. From a conve…

2022

Efficient Model-based Multi-agent Reinforcement Learning via Optimistic Equilibrium Computation

ICML 2022spotlight

We consider model-based multi-agent reinforcement learning, where the environment transition model is unknown and can only be learned via expensive interactions with the environment. We propose H-MARL (Hallucinated Multi-Agent Reinforcement Learning), a novel sample-efficient algorithm that can effi…

Cited by 19SourcePDFScholar
2021

Online Submodular Resource Allocation with Applications to Rebalancing Shared Mobility Systems

ICML 2021spotlight

Motivated by applications in shared mobility, we address the problem of allocating a group of agents to a set of resources to maximize a cumulative welfare objective. We model the welfare obtainable from each resource as a monotone DR-submodular function which is a-priori unknown and can only be lea…

Cited by 3SourcePDFScholar
2020

Contextual Games: Multi-Agent Learning with Side Information

NeurIPS 2020poster

We formulate the novel class of contextual games, a type of repeated games driven by contextual information at each round. By means of kernel-based regularity assumptions, we model the correlation between different contexts and game outcomes and propose a novel online (meta) algorithm that exploits…

Cited by 23SourcePDFScholar
2020

Learning to Play Sequential Games versus Unknown Opponents

NeurIPS 2020poster

We consider a repeated sequential game between a learner, who plays first, and an opponent who responds to the chosen action. We seek to design strategies for the learner to successfully interact with the opponent. While most previous approaches consider known opponent models, we focus on the settin…

Cited by 27SourcePDFScholar
2020

Mixed Strategies for Robust Optimization of Unknown Objectives

AISTATS 2020poster

We consider robust optimization problems, where the goal is to optimize an unknown objective function against the worst-case realization of an uncertain parameter. For this setting, we design a novel sample-efficient algorithm GP-MRO, which sequentially learns about the unknown objective from noisy…

Cited by 18SourcePDFScholar
2019

Bounding Inefficiency of Equilibria in Continuous Actions Games using Submodularity and Curvature

AISTATS 2019poster

Games with continuous strategy sets arise in several machine learning problems (e.g. adversarial learning). For such games, simple no-regret learning algorithms exist in several cases and ensure convergence to coarse correlated equilibria (CCE). The efficiency of such equilibria with respect to a s…

Cited by 14SourcePDFScholar
2019

No-Regret Learning in Unknown Games with Correlated Payoffs

NeurIPS 2019poster

We consider the problem of learning to play a repeated multi-agent game with an unknown reward function. Single player online learning algorithms attain strong regret bounds when provided with full information feedback, which unfortunately is unavailable in many real-world scenarios. Bandit feedback…

Cited by 48SourcePDFScholar