← Search

Shuze Liu

6 accepted papers

2026

Convergence of Two-Timescale Stochastic Approximation with Markovian Samples and Applications in Reinforcement Learning

ICML 2026poster

Stochastic approximations (SA)--algorithms which derive their power through the use of random, incremental updates--are at the heart of reinforcement learning (RL). Expanding the theory of SA has established rigorous results concerning the most important algorithms in RL, including stochastic gradie…

Cited by 0SourceScholar
2026

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

ICML 2026poster

While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models (LLMs), the absence of many folklore lemmas in Mathlib remains a persistent barrier that limits Lean's usability as an everyday tool for mathematicians like …

Cited by 0SourceScholar
2026

Offline Two-Player Zero-Sum Markov Games with KL Regularization

ICML 2026poster

We study the problem of learning Nash equilibria in offline two-player zero-sum Markov games. While existing approaches often rely on explicit pessimism to address distribution shift, we show that KL regularization alone suffices to stabilize learning and guarantee convergence. We first introduce Re…

Cited by 0SourceScholar
2025

Efficient Policy Evaluation with Safety Constraint for Reinforcement Learning

ICLR 2025poster

In reinforcement learning, classic on-policy evaluation methods often suffer from high variance and require massive online data to attain the desired accuracy. Previous studies attempt to reduce evaluation variance by searching for or designing proper behavior policies to collect data. However, thes…

Cited by 3SourcePDFScholar