← Search

Zuyuan Zhang

5 accepted papers

2026

HodgeFlow Policy Search by Topologically Dissecting Temporal-Difference Signals in Non-Markovian Environments

ICML 2026poster

Non-Markovian dynamics are commonly found in real-world environments due to long-range dependencies, partial observability, and memory effects. The Bellman equation that is the central pillar of Reinforcement learning (RL) becomes only approximately valid under Non-Markovian. Existing work often foc…

Cited by 0SourceScholar
2026

NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search

ICML 2026spotlight

Monte Carlo Tree Search (MCTS) scales poorly in cooperative multi-agent domains because expansion must consider an exponentially large set of joint actions, severely limiting exploration under realistic search budgets. We propose \textsc{NonZero}, which keeps multi-agent MCTS tractable by running su…

Cited by 0SourceScholar
2025

Learning to Collaborate with Unknown Agents in the Absence of Reward

AAAI 2025technical

With the advancements of artificial intelligence (AI), emerging scenarios involving close collaboration between AI and other unknown agents are becoming increasingly common. This requires sometimes training AI agents to collaborate with unknown agents in the absence of a reward function -- which may…

Cited by 0SourcePDFScholar
2025

Second-Order Convergence in Private Stochastic Non-Convex Optimization

NeurIPS 2025poster

We investigate the problem of finding second-order stationary points (SOSP) in differentially private (DP) stochastic non-convex optimization. Existing methods suffer from two key limitations: \textbf{(i)} inaccurate convergence error rate due to overlooking gradient variance in the saddle point esc…

Cited by 0SourceScholar