← Search

Zhaoyi Zhou

4 accepted papers

2026

Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards

ICML 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful paradigm for post-training large reasoning models (LRMs) using policy-gradient methods such as GRPO. To stabilize training, these methods typically center trajectory rewards by subtracting the empirical mean for each pro…

Cited by 0SourceScholar
2024

Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning

ICLR 2024poster

Off-policy dynamic programming (DP) techniques such as $Q$-learning have proven to be important in sequential decision-making problems. In the presence of function approximation, however, these techniques often diverge due to the absence of Bellman completeness in the function classes considered, a…

2023

Convergence rates for localized actor-critic in networked Markov potential games

UAI 2023poster

We introduce a class of networked Markov potential games where agents are associated with nodes in a network. Each agent has its own local potential function, and the reward of each agent depends only on the states and actions of agents within a neighborhood. In this context, we propose a localized…