← Search

Kaiqing Zhang

39 accepted papers

2026

Online Learning and Equilibrium Computation with Ranking Feedback

ICLR 2026oral

Online learning in arbitrary and possibly adversarial environments has been extensively studied in sequential decision-making, with a strong connection to equilibrium computation in game theory. Most existing online learning algorithms are based on \emph{numeric} utility feedback from the environmen…

Cited by 0SourceScholar
2026

Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach

ICML 2026poster

Large language models (LLMs) are increasingly deployed as agents for decision-making (DM) in interactive and dynamic environments. However, since they are not originally designed for DM, recent studies show that LLMs struggle in basic online DM settings. We introduce ITERATIVE REGRET-MINIMIZATION FI…

Cited by 0SourceScholar
2025

Do LLM Agents Have Regret? A Case Study in Online Learning and Games

ICLR 2025poster

Large language models (LLMs) have been increasingly employed for (interactive) decision-making, via the development of LLM-based autonomous agents. Despite their emerging successes, the performance of LLM agents in decision-making has not been fully investigated through quantitative metrics, especia…

Cited by 20SourcePDFScholar
2025

Foundations of Multi-Agent Learning in Dynamic Environments: Where Reinforcement Learning Meets Strategic Decision-Making

AAAI 2025technical

Recent years have witnessed tremendous successes of learning for sequential decision-making, and in particular, Reinforcement Learning (RL). Prominent application examples include playing Go and video games, robotics, autonomous driving, and recently large language models. Most such success stories…

Cited by 0SourcePDFScholar
2025

MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning

ACL 2025long

Leveraging multi-agentic frameworks to enhance large language models (LLMs) has demonstrated significant potential recently, with most existing studies focusing on prompting and developing workflows with frozen LLMs. In this paper, we aim to further unleash the power of such multi-agentic frameworks…

Cited by 0SourcePDFScholar
2024

Provable Partially Observable Reinforcement Learning with Privileged Information

NeurIPS 2024poster

Partial observability of the underlying states generally presents significant challenges for reinforcement learning (RL). In practice, certain *privileged information* , e.g., the access to states from simulators, has been exploited in training and achieved prominent empirical successes. To better u…

Cited by 2SourcePDFScholar
2024

Robot Fleet Learning via Policy Merging

ICLR 2024poster

Fleets of robots ingest massive amounts of heterogeneous streaming data silos generated by interacting with their environments, far more than what can be stored or transmitted with ease. At the same time, teams of robots should co-acquire diverse skills through their heterogeneous experiences in var…

2023

A Finite-Sample Analysis of Payoff-Based Independent Learning in Zero-Sum Stochastic Games

NeurIPS 2023poster

In this work, we study two-player zero-sum stochastic games and develop a variant of the smoothed best-response learning dynamics that combines independent learning dynamics for matrix games with the minimax value iteration for stochastic games. The resulting learning dynamics are payoff-based, conv…

Cited by 14SourcePDFScholar
2023

Byzantine-Robust Online and Offline Distributed Reinforcement Learning

AISTATS 2023poster

We consider a distributed reinforcement learning setting where multiple agents separately explore the environment and communicate their experiences through a central server. However, $\alpha$-fraction of agents are adversarial and can report arbitrary fake information. Critically, these adversarial…

Cited by 23SourcePDFScholar
2023

Does Learning from Decentralized Non-IID Unlabeled Data Benefit from Self Supervision?

ICLR 2023poster

The success of machine learning relies heavily on massive amounts of data, which are usually generated and stored across a range of diverse and distributed data sources. Decentralized learning has thus been advocated and widely deployed to make efficient use of distributed datasets, with an extensiv…

2023

Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs

NeurIPS 2023poster

We study the problem of computing an optimal policy of an infinite-horizon discounted constrained Markov decision process (constrained MDP). Despite the popularity of Lagrangian-based policy search methods used in practice, the oscillation of policy iterates in these methods has not been fully under…

Cited by 30SourcePDFScholar
2023

Learning to Extrapolate: A Transductive Approach

ICLR 2023poster

Machine learning systems, especially with overparameterized deep neural networks, can generalize to novel test instances drawn from the same distribution as the training data. However, they fare poorly when evaluated on out-of-support test points. In this work, we tackle the problem of developing ma…

2023

Multi-Player Zero-Sum Markov Games with Networked Separable Interactions

NeurIPS 2023poster

We study a new class of Markov games, \textit{(multi-player) zero-sum Markov Games} with {\it Networked separable interactions} (zero-sum NMGs), to model the local interaction structure in non-cooperative multi-agent sequential decision-making. We define a zero-sum NMG as a model where {the payoffs…

Cited by 11SourcePDFScholar
2023

Partially Observable Multi-agent RL with (Quasi-)Efficiency: The Blessing of Information Sharing

ICML 2023poster

We study provable multi-agent reinforcement learning (MARL) in the general framework of partially observable stochastic games (POSGs). To circumvent the known hardness results and the use of computationally intractable oracles, we propose to leverage the potential *information-sharing* among agents,…

Cited by 10SourcePDFScholar
2023

Revisiting the Linear-Programming Framework for Offline RL with General Function Approximation

ICML 2023poster

Offline reinforcement learning (RL) aims to find an optimal policy for sequential decision-making using a pre-collected dataset, without further interaction with the environment. Recent theoretical progress has focused on developing sample-efficient offline RL algorithms with various relaxed assumpt…

Cited by 27SourcePDFScholar
2023

Self-Supervised Reinforcement Learning that Transfers using Random Features

NeurIPS 2023poster

Model-free reinforcement learning algorithms have exhibited great potential in solving single-task sequential decision-making problems with high-dimensional observations and long horizons, but are known to be hard to generalize across tasks. Model-based RL, on the other hand, learns task-agnostic mo…

Cited by 11SourcePDFScholar
2023

Symmetric (Optimistic) Natural Policy Gradient for Multi-Agent Learning with Parameter Convergence

AISTATS 2023poster

Multi-agent interactions are increasingly important in the context of reinforcement learning, and the theoretical foundations of policy gradient methods have attracted surging research interest. We investigate the global convergence of natural policy gradient (NPG) algorithms in multi-agent learning…

Cited by 15SourcePDFScholar
2023

The Power of Regularization in Solving Extensive-Form Games

ICLR 2023poster

In this paper, we investigate the power of {\it regularization}, a common technique in reinforcement learning and optimization, in solving extensive-form games (EFGs). We propose a series of new algorithms based on regularizing the payoff functions of the game, and establish a set of convergence re…

Cited by 26SourcePDFScholar
2022

Do Differentiable Simulators Give Better Policy Gradients?

ICML 2022oral

Differentiable simulators promise faster computation time for reinforcement learning by replacing zeroth-order gradient estimates of a stochastic objective with an estimate based on first-order gradients. However, it is yet unclear what factors decide the performance of the two estimators on complex…

Cited by 127SourcePDFScholar
2022

Globally Convergent Policy Search for Output Estimation

NeurIPS 2022accept

We introduce the first direct policy search algorithm which provably converges to the globally optimal dynamic filter for the classical problem of predicting the outputs of a linear dynamical system, given noisy, partial observations. Despite the ubiquity of partial observability in practice, theore…

Cited by 14SourcePDFScholar
2022

Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic Convergence

ICML 2022oral

We examine global non-asymptotic convergence properties of policy gradient methods for multi-agent reinforcement learning (RL) problems in Markov potential games (MPGs). To learn a Nash equilibrium of an MPG in which the size of state space and/or the number of players can be very large, we propose…

Cited by 97SourcePDFScholar
2022

What is a Good Metric to Study Generalization of Minimax Learners?

NeurIPS 2022accept

Minimax optimization has served as the backbone of many machine learning problems. Although the convergence behavior of optimization algorithms has been extensively studied in minimax settings, their generalization guarantees, i.e., how the model trained on empirical data performs on the unseen test…

Cited by 16SourcePDFScholar
2021

Decentralized Policy Gradient Descent Ascent for Safe Multi-Agent Reinforcement Learning

AAAI 2021technical

This paper deals with distributed reinforcement learning problems with safety constraints. In particular, we consider that a team of agents cooperate in a shared environment, where each agent has its individual reward function and safety constraints that involve all agents' joint actions. As such, t…

Cited by 79SourcePDFScholar
2021

Decentralized Q-learning in Zero-sum Markov Games

NeurIPS 2021poster

We study multi-agent reinforcement learning (MARL) in infinite-horizon discounted zero-sum Markov games. We focus on the practical but challenging setting of decentralized MARL, where agents make decisions without coordination by a centralized controller, but only based on their own payoffs and lo…

Cited by 121SourcePDFScholar
2021

Derivative-Free Policy Optimization for Linear Risk-Sensitive and Robust Control Design: Implicit Regularization and Sample Complexity

NeurIPS 2021poster

Direct policy search serves as one of the workhorses in modern reinforcement learning (RL), and its applications in continuous control tasks have recently attracted increasing attention. In this work, we investigate the convergence theory of policy gradient (PG) methods for learning the linear risk-…

Cited by 62SourcePDFScholar
2021

Learning Safe Multi-agent Control with Decentralized Neural Barrier Certificates

ICLR 2021poster

We study the multi-agent safe control problem where agents should avoid collisions to static obstacles and collisions with each other while reaching their goals. Our core idea is to learn the multi-agent control policy jointly with learning the control barrier functions as safety certificates. We p…

Cited by 175SourcePDFScholar
2021

Near-Optimal Model-Free Reinforcement Learning in Non-Stationary Episodic MDPs

ICML 2021spotlight

We consider model-free reinforcement learning (RL) in non-stationary Markov decision processes. Both the reward functions and the state transition functions are allowed to vary arbitrarily over time as long as their cumulative variations do not exceed certain variation budgets. We propose Restarted…

Cited by 49SourcePDFScholar
2021

Reinforcement Learning for Cost-Aware Markov Decision Processes

ICML 2021spotlight

Ratio maximization has applications in areas as diverse as finance, reward shaping for reinforcement learning (RL), and the development of safe artificial intelligence, yet there has been very little exploration of RL algorithms for ratio maximization. This paper addresses this deficiency by introdu…

Cited by 11SourcePDFScholar
2020

An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient Methods

NeurIPS 2020poster

In this paper, we revisit and improve the convergence of policy gradient (PG), natural PG (NPG) methods, and their variance-reduced variants, under general smooth policy parametrizations. More specifically, with the Fisher information matrix of the policy being positive definite: i) we show that a s…

2020

Model-Based Multi-Agent RL in Zero-Sum Markov Games with Near-Optimal Sample Complexity

NeurIPS 2020spotlight

Model-based reinforcement learning (RL), which finds an optimal policy using an empirical model, has long been recognized as one of the cornerstones of RL. It is especially suitable for multi-agent RL (MARL), as it naturally decouples the learning and the planning phases, and avoids the non-stationa…

Cited by 169SourcePDFScholar
2020

Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision Processes

NeurIPS 2020poster

We study sequential decision-making problems in which each agent aims to maximize the expected total reward while satisfying a constraint on the expected total utility. We employ the natural policy gradient method to solve the discounted infinite-horizon Constrained Markov Decision Processes (CMDPs)…

Cited by 241SourcePDFScholar
2020

On the Stability and Convergence of Robust Adversarial Reinforcement Learning: A Case Study on Linear Quadratic Systems

NeurIPS 2020poster

Reinforcement learning (RL) algorithms can fail to generalize due to the gap between the simulation and the real world. One standard remedy is to use robust adversarial RL (RARL) that accounts for this gap during the policy training, by modeling the gap as an adversary against the training agent. In…

Cited by 61SourcePDFScholar
2020

POLY-HOOT: Monte-Carlo Planning in Continuous Space MDPs with Non-Asymptotic Analysis

NeurIPS 2020poster

Monte-Carlo planning, as exemplified by Monte-Carlo Tree Search (MCTS), has demonstrated remarkable performance in applications with finite spaces. In this paper, we consider Monte-Carlo planning in an environment with continuous state-action spaces, a much less understood problem with important app…

Cited by 24SourcePDFScholar
2020

Robust Multi-Agent Reinforcement Learning with Model Uncertainty

NeurIPS 2020poster

In this work, we study the problem of multi-agent reinforcement learning (MARL) with model uncertainty, which is referred to as robust MARL. This is naturally motivated by some multi-agent applications where each agent may not have perfectly accurate knowledge of the model, e.g., all the reward func…

Cited by 111SourcePDFScholar
2019

Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

NeurIPS 2019poster

We study the global convergence of policy optimization for finding the Nash equilibria (NE) in zero-sum linear quadratic (LQ) games. To this end, we first investigate the landscape of LQ games, viewing it as a nonconvex-nonconcave saddle-point problem in the policy space. Specifically, we show that…

Cited by 161SourcePDFScholar
2018

Fully Decentralized Multi-Agent Reinforcement Learning with Networked Agents

ICML 2018oral

We consider the fully decentralized multi-agent reinforcement learning (MARL) problem, where the agents are connected via a time-varying and possibly sparse communication network. Specifically, we assume that the reward functions of the agents might correspond to different tasks, and are only known…

Cited by 786SourcePDFScholar
2018

Nonlinear Structured Signal Estimation in High Dimensions via Iterative Hard Thresholding

AISTATS 2018poster

We study the high-dimensional signal estimation problem with nonlinear measurements, where the signal of interest is either sparse or low-rank. In both settings, our estimator is formulated as the minimizer of the nonlinear least-squares loss function under a combinatorial constraint, which is obtai…

Cited by 0SourcePDFScholar