← Search

Canzhe Zhao

7 accepted papers

2025

Learning Imperfect Information Extensive-form Games with Last-iterate Convergence under Bandit Feedback

ICML 2025poster

We investigate learning approximate Nash equilibrium (NE) policy profiles in two-player zero-sum imperfect information extensive-form games (IIEFGs) with last-iterate convergence guarantees. Existing algorithms either rely on full-information feedback or provide only asymptotic convergence rates. In…

Cited by 0SourcePDFScholar
2025

Logarithmic Regret for Linear Markov Decision Processes with Adversarial Corruptions

AAAI 2025technical

In this work, we study the logarithmic regret for reinforcement learning (RL) with linear function approximation and adversarial corruptions, in the formulation of linear Markov decision processes (MDPs). Specifically, we consider the case where there exist adversarial corruptions over the reward fu…

Cited by 0SourcePDFScholar
2025

Towards Provably Efficient Learning of Imperfect Information Extensive-Form Games with Linear Function Approximation

UAI 2025

Despite significant advances in learning imperfect information extensive-form games (IIEFGs), most existing theoretical guarantees are limited to IIEFGs in the tabular case. To permit efficient learning of large-scale IIEFGs, we take the first step in studying two-player zero-sum IIEFGs with linear

2023

DPMAC: Differentially Private Communication for Cooperative Multi-Agent Reinforcement Learning

IJCAI 2023poster

Communication lays the foundation for cooperation in human society and in multi-agent reinforcement learning (MARL). Humans also desire to maintain their privacy when communicating with others, yet such privacy concern has not been considered in existing works in MARL. We propose the differentially…

2023

Learning Adversarial Linear Mixture Markov Decision Processes with Bandit Feedback and Unknown Transition

ICLR 2023poster

We study reinforcement learning (RL) with linear function approximation, unknown transition, and adversarial losses in the bandit feedback setting. Specifically, the unknown transition probability function is a linear mixture model \citep{AyoubJSWY20,ZhouGS21,HeZG22} with a given feature mapping, an…

Cited by 13SourcePDFScholar
2023

Learning Adversarial Low-rank Markov Decision Processes with Unknown Transition and Full-information Feedback

NeurIPS 2023poster

In this work, we study the low-rank MDPs with adversarially changed losses in the full-information feedback setting. In particular, the unknown transition probability kernel admits a low-rank matrix decomposition \citep{REPUCB22}, and the loss functions may change adversarially but are revealed to t…

Cited by 5SourcePDFScholar
2022

Simultaneously Learning Stochastic and Adversarial Bandits under the Position-Based Model

AAAI 2022technical

Online learning to rank (OLTR) interactively learns to choose lists of items from a large collection based on certain click models that describe users' click behaviors. Most recent works for this problem focus on the stochastic environment where the item attractiveness is assumed to be invariant dur…

Cited by 6SourcePDFScholar