← Search

Zhihan Xiong

8 accepted papers

2024

A Black-box Approach for Non-stationary Multi-agent Reinforcement Learning

ICLR 2024poster

We investigate learning the equilibria in non-stationary multi-agent systems and address the challenges that differentiate multi-agent learning from single-agent learning. Specifically, we focus on games with bandit feedback, where testing an equilibrium can result in substantial regret even when th…

Cited by 2SourcePDFScholar
2024

A/B Testing and Best-arm Identification for Linear Bandits with Robustness to Non-stationarity

AISTATS 2024poster

We investigate the fixed-budget best-arm identification (BAI) problem for linear bandits in a potentially non-stationary environment. Given a finite arm set $\mathcal{X}\subset\mathbb{R}^d$, a fixed budget $T$, and an unpredictable sequence of parameters $\left\lbrace\theta_t\right\rbrace_{t=1}^{T}$…

2023

Offline Congestion Games: How Feedback Type Affects Data Coverage Requirement

ICLR 2023poster

This paper investigates when one can efficiently recover an approximate Nash Equilibrium (NE) in offline congestion games. The existing dataset coverage assumption in offline general-sum games inevitably incurs a dependency on the number of actions, which can be exponentially large in congestion gam…

Cited by 1SourcePDFScholar
2022

Near-Optimal Randomized Exploration for Tabular Markov Decision Processes

NeurIPS 2022accept

We study algorithms using randomized value functions for exploration in reinforcement learning. This type of algorithms enjoys appealing empirical performance. We show that when we use 1) a single random seed in each episode, and 2) a Bernstein-type magnitude of noise, we obtain a worst-case $\widet…

Cited by 10SourcePDFScholar
2021

Selective Sampling for Online Best-arm Identification

NeurIPS 2021poster

This work considers the problem of selective-sampling for best-arm identification. Given a set of potential options $\mathcal{Z}\subset\mathbb{R}^d$, a learner aims to compute with probability greater than $1-\delta$, $\arg\max_{z\in \mathcal{Z}} z^{\top}\theta_{\ast}$ where $\theta_{\ast}$ is unkno…

Cited by 8SourcePDFScholar