← Search

Sizhe Tang

3 accepted papers

2026

HodgeFlow Policy Search by Topologically Dissecting Temporal-Difference Signals in Non-Markovian Environments

ICML 2026poster

Non-Markovian dynamics are commonly found in real-world environments due to long-range dependencies, partial observability, and memory effects. The Bellman equation that is the central pillar of Reinforcement learning (RL) becomes only approximately valid under Non-Markovian. Existing work often foc…

Cited by 0SourceScholar
2026

NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search

ICML 2026spotlight

Monte Carlo Tree Search (MCTS) scales poorly in cooperative multi-agent domains because expansion must consider an exponentially large set of joint actions, severely limiting exploration under realistic search budgets. We propose \textsc{NonZero}, which keeps multi-agent MCTS tractable by running su…

Cited by 0SourceScholar
2025

MALinZero: Efficient Low-Dimensional Search for Mastering Complex Multi-Agent Planning

NeurIPS 2025poster

Monte Carlo Tree Search (MCTS), which leverages Upper Confidence Bound for Trees (UCTs) to balance exploration and exploitation through randomized sampling, is instrumental to solving complex planning problems. However, for multi-agent planning, MCTS is confronted with a large combinatorial action s…

Cited by 0SourceScholar