← Search

YUANHAO WANG

14 accepted papers

2025

Securing Equal Share: A Principled Approach for Learning Multiplayer Symmetric Games

ICML 2025poster

This paper examines multiplayer symmetric constant-sum games with more than two players in a competitive setting, such as Mahjong, Poker, and various board and video games. In contrast to two-player zero-sum games, equilibria in multiplayer games are neither unique nor non-exploitable, failing to pr…

Cited by 0SourcePDFScholar
2024

Directional Smoothness and Gradient Methods: Convergence and Adaptivity

NeurIPS 2024poster

We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization, rather than on global, worst-case constants. Key to our proofs is directional smoothness, a measure of gradient variation that we use to develop upper-boun…

Cited by 5SourcePDFScholar
2023

Learning Adaptive Tensorial Density Fields for Clean Cryo-ET Reconstruction

NeurIPS 2023poster

We present a novel learning-based framework for reconstructing 3D structures from tilt-series cryo-Electron Tomography (cryo-ET) data. Cryo-ET is a powerful imaging technique that can achieve near-atomic resolutions. Still, it suffers from challenges such as missing-wedge acquisition, large data siz…

2022

Learning Markov Games with Adversarial Opponents: Efficient Algorithms and Fundamental Limits

ICML 2022oral

An ideal strategy in zero-sum games should not only grant the player an average reward no less than the value of Nash equilibrium, but also exploit the (adaptive) opponents when they are suboptimal. While most existing works in Markov games focus exclusively on the former objective, it remains open…

Cited by 25SourcePDFScholar
2022

Near-optimal Local Convergence of Alternating Gradient Descent-Ascent for Minimax Optimization

AISTATS 2022poster

Smooth minimax games often proceed by simultaneous or alternating gradient updates. Although algorithms with alternating updates are commonly used in practice, the majority of existing theoretical analyses focus on simultaneous algorithms for convenience of analysis. In this paper, we study alternat…

Cited by 59SourcePDFScholar
2021

An Exponential Lower Bound for Linearly Realizable MDP with Constant Suboptimality Gap

NeurIPS 2021oral

A fundamental question in the theory of reinforcement learning is: suppose the optimal $Q$-function lies in the linear span of a given $d$ dimensional feature mapping, is sample-efficient reinforcement learning (RL) possible? The recent and remarkable result of Weisz et al. (2020) resolves this ques…

Cited by 58SourcePDFScholar
2020

Distributed Bandit Learning: Near-Optimal Regret with Efficient Communication

ICLR 2020poster

We study the problem of regret minimization for distributed bandits learning, in which $M$ agents work collaboratively to minimize their total regret under the coordination of a central server. Our goal is to design communication protocols with near-optimal regret and little communication cost, whic…

Cited by 105SourceScholar
2020

Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP

ICLR 2020poster

A fundamental question in reinforcement learning is whether model-free algorithms are sample efficient. Recently, Jin et al. (2018) proposed a Q-learning algorithm with UCB exploration policy, and proved it has nearly optimal regret bound for finite-horizon episodic MDP. In this paper, we adapt Q-l…

Cited by 125SourceScholar
2020

Stereo Event-based Particle Tracking Velocimetry for 3D Fluid Flow Reconstruction

ECCV 2020poster

Existing Particle Imaging Velocimetry techniques require the use of high-speed cameras to reconstruct time-resolved fluid flows. These cameras provides high-resolution images at high frame rates, which generates bandwidth and memory issues. By capturing only changes in the brightness with a very low…