← Search

Tamer Basar

17 accepted papers

2023

Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task Similarity

NeurIPS 2023poster

Multi-agent reinforcement learning (MARL) has primarily focused on solving a single task in isolation, while in practice the environment is often evolving, leaving many related tasks to be solved. In this paper, we investigate the benefits of meta-learning in solving multiple MARL tasks collectively…

Cited by 8SourcePDFScholar
2023

Oracle-free Reinforcement Learning in Mean-Field Games along a Single Sample Path

AISTATS 2023poster

We consider online reinforcement learning in Mean-Field Games (MFGs). Unlike traditional approaches, we alleviate the need for a mean-field oracle by developing an algorithm that approximates the Mean-Field Equilibrium (MFE) using the single sample path of the generic agent. We call this Sandbox Lea…

Cited by 31SourcePDFScholar
2022

A Mean-Field Game Approach to Cloud Resource Management with Function Approximation

NeurIPS 2022accept

Reinforcement learning (RL) has gained increasing popularity for resource management in cloud services such as serverless computing. As self-interested users compete for shared resources in a cluster, the multi-tenancy nature of serverless platforms necessitates multi-agent reinforcement learning (M…

Cited by 26SourcePDFScholar
2021

Decentralized Q-learning in Zero-sum Markov Games

NeurIPS 2021poster

We study multi-agent reinforcement learning (MARL) in infinite-horizon discounted zero-sum Markov games. We focus on the practical but challenging setting of decentralized MARL, where agents make decisions without coordination by a centralized controller, but only based on their own payoffs and lo…

Cited by 121SourcePDFScholar
2021

Derivative-Free Policy Optimization for Linear Risk-Sensitive and Robust Control Design: Implicit Regularization and Sample Complexity

NeurIPS 2021poster

Direct policy search serves as one of the workhorses in modern reinforcement learning (RL), and its applications in continuous control tasks have recently attracted increasing attention. In this work, we investigate the convergence theory of policy gradient (PG) methods for learning the linear risk-…

Cited by 62SourcePDFScholar
2021

Near-Optimal Model-Free Reinforcement Learning in Non-Stationary Episodic MDPs

ICML 2021spotlight

We consider model-free reinforcement learning (RL) in non-stationary Markov decision processes. Both the reward functions and the state transition functions are allowed to vary arbitrarily over time as long as their cumulative variations do not exceed certain variation budgets. We propose Restarted…

Cited by 49SourcePDFScholar
2020

An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient Methods

NeurIPS 2020poster

In this paper, we revisit and improve the convergence of policy gradient (PG), natural PG (NPG) methods, and their variance-reduced variants, under general smooth policy parametrizations. More specifically, with the Fisher information matrix of the policy being positive definite: i) we show that a s…

2020

Model-Based Multi-Agent RL in Zero-Sum Markov Games with Near-Optimal Sample Complexity

NeurIPS 2020spotlight

Model-based reinforcement learning (RL), which finds an optimal policy using an empirical model, has long been recognized as one of the cornerstones of RL. It is especially suitable for multi-agent RL (MARL), as it naturally decouples the learning and the planning phases, and avoids the non-stationa…

Cited by 169SourcePDFScholar
2020

Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision Processes

NeurIPS 2020poster

We study sequential decision-making problems in which each agent aims to maximize the expected total reward while satisfying a constraint on the expected total utility. We employ the natural policy gradient method to solve the discounted infinite-horizon Constrained Markov Decision Processes (CMDPs)…

Cited by 241SourcePDFScholar
2020

On the Stability and Convergence of Robust Adversarial Reinforcement Learning: A Case Study on Linear Quadratic Systems

NeurIPS 2020poster

Reinforcement learning (RL) algorithms can fail to generalize due to the gap between the simulation and the real world. One standard remedy is to use robust adversarial RL (RARL) that accounts for this gap during the policy training, by modeling the gap as an adversary against the training agent. In…

Cited by 61SourcePDFScholar
2020

POLY-HOOT: Monte-Carlo Planning in Continuous Space MDPs with Non-Asymptotic Analysis

NeurIPS 2020poster

Monte-Carlo planning, as exemplified by Monte-Carlo Tree Search (MCTS), has demonstrated remarkable performance in applications with finite spaces. In this paper, we consider Monte-Carlo planning in an environment with continuous state-action spaces, a much less understood problem with important app…

Cited by 24SourcePDFScholar
2020

Robust Multi-Agent Reinforcement Learning with Model Uncertainty

NeurIPS 2020poster

In this work, we study the problem of multi-agent reinforcement learning (MARL) with model uncertainty, which is referred to as robust MARL. This is naturally motivated by some multi-agent applications where each agent may not have perfectly accurate knowledge of the model, e.g., all the reward func…

Cited by 111SourcePDFScholar
2019

Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

NeurIPS 2019poster

We study the global convergence of policy optimization for finding the Nash equilibria (NE) in zero-sum linear quadratic (LQ) games. To this end, we first investigate the landscape of LQ games, viewing it as a nonconvex-nonconcave saddle-point problem in the policy space. Specifically, we show that…

Cited by 161SourcePDFScholar
2018

Fully Decentralized Multi-Agent Reinforcement Learning with Networked Agents

ICML 2018oral

We consider the fully decentralized multi-agent reinforcement learning (MARL) problem, where the agents are connected via a time-varying and possibly sparse communication network. Specifically, we assume that the reward functions of the agents might correspond to different tasks, and are only known…

Cited by 786SourcePDFScholar