← Search

Alexandre Bayen

7 accepted papers

2022

The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games

NeurIPS 2022accept

Proximal Policy Optimization (PPO) is a ubiquitous on-policy reinforcement learning algorithm but is significantly less utilized than off-policy learning algorithms in multi-agent settings. This is often due to the belief that PPO is significantly less sample efficient than off-policy methods in mul…

2020

Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design

NeurIPS 2020oral

A wide range of reinforcement learning (RL) problems --- including robustness, transfer learning, unsupervised RL, and emergent complexity --- require specifying a distribution of tasks or environments in which a policy will be trained. However, creating a useful distribution of environments is err…

2017

Random projection design for scalable implicit smoothing of randomly observed stochastic processes

AISTATS 2017poster

Sampling at random timestamps, long range dependencies, and scale hamper standard meth- ods for multivariate time series analysis. In this paper we present a novel estimator for cross-covariance of randomly observed time series which unravels the dynamics of an unobserved stochastic process. We anal…

Cited by 8SourcePDFScholar
2016

Minimizing Regret on Reflexive Banach Spaces and Nash Equilibria in Continuous Zero-Sum Games

NeurIPS 2016poster

We study a general adversarial online learning problem, in which we are given a decision set X' in a reflexive Banach space X and a sequence of reward vectors in the dual space of X. At each iteration, we choose an action from X', based on the observed sequence of previous rewards. Our goal is to mi…

Cited by 17SourcePDFScholar
2015

Accelerated Mirror Descent in Continuous and Discrete Time

NeurIPS 2015spotlight

We study accelerated mirror descent dynamics in continuous and discrete time. Combining the original continuous-time motivation of mirror descent with a recent ODE interpretation of Nesterov's accelerated method, we propose a family of continuous-time descent dynamics for convex functions with Lipsc…

Cited by 320SourcePDFScholar