← Search

Zongkai Liu

6 accepted papers

2026

DisPPO: Quantile-Based Distributional Reinforcement Learning for Large Language Models

ICML 2026poster

Reinforcement Learning (RL) has become a cornerstone for enhancing the reasoning capabilities of Large Language Models (LLMs). However, standard actor-critic methods, such as PPO, rely on scalar value functions that estimate only the expectation of cumulative returns. This reduction inherently disca…

Cited by 0SourceScholar
2025

Conservative Offline Goal-Conditioned Implicit V-Learning

ICML 2025poster

Offline goal-conditioned reinforcement learning (GCRL) learns a goal-conditioned value function to train policies for diverse goals with pre-collected datasets. Hindsight experience replay addresses the issue of sparse rewards by treating intermediate states as goals but fails to complete goal-stitc…

Cited by 0SourcePDFScholar
2025

Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization

AAAI 2025technical

Offline Multi-Agent Reinforcement Learning (MARL) is an emerging field that aims to learn optimal multi-agent policies from pre-collected datasets. Compared to single-agent case, multi-agent setting involves a large joint state-action space and coupled behaviors of multiple agents, which bring extra…

2025

Rapid Learning in Constrained Minimax Games with Negative Momentum

AAAI 2025technical

In this paper, we delve into the utilization of the negative momentum technique in constrained minimax games. From an intuitive mechanical standpoint, we introduce a novel framework for momentum buffer updating, which extends the findings of negative momentum from the unconstrained setting to the co…

2024

An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement Learning

NeurIPS 2024poster

In recent years, significant progress has been made in multi-objective reinforcement learning (RL) research, which aims to balance multiple objectives by incorporating preferences for each objective. In most existing studies, specific preferences must be provided during deployment to indicate the de…

Cited by 0SourcePDFScholar
2022

A Unified Diversity Measure for Multiagent Reinforcement Learning

NeurIPS 2022accept

Promoting behavioural diversity is of critical importance in multi-agent reinforcement learning, since it helps the agent population maintain robust performance when encountering unfamiliar opponents at test time, or, when the game is highly non-transitive in the strategy space (e.g., Rock-Paper-Sc…

Cited by 16SourcePDFScholar