← Search

Qianchuan Zhao

20 accepted papers

2026

From Winning to Understanding: A Diagnostic Long-Horizon RTS Benchmark for LLMs

ICML 2026poster

Large language models (LLMs) are increasingly used as decision modules, yet existing benchmarks provide limited coverage of long-horizon, adversarial interaction while faithfully acting on human instructions. We introduce a long-horizon Red Alert RTS benchmark with a hierarchical interface in which …

Cited by 0SourceScholar
2026

GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent System

ICLR 2026poster

In the realm of multi-agent systems, the challenge of partial observability is a critical barrier to effective coordination and decision-making. Existing approaches, such as belief state estimation and inter-agent communication, often fall short. Belief-based methods are limited by their focus on pa…

Cited by 0SourceScholar
2026

Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language Models

ICLR 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for enhancing Large Language Models (LLMs) on complex reasoning tasks. Yet current methods face an exploration dilemma: standard RL struggles to escape the local optima of pre-trained LLMs’ sharply peaked initial policies, bo…

Cited by 0SourceScholar
2025

DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning

NeurIPS 2025poster

Comprehending natural language and following human instructions are critical capabilities for intelligent agents. However, the flexibility of linguistic instructions induces substantial ambiguity across language-conditioned tasks, severely degrading algorithmic performance. To address these limitat…

Cited by 0SourcecodeScholar
2025

Episodic Novelty Through Temporal Distance

ICLR 2025poster

Exploration in sparse reward environments remains a significant challenge in reinforcement learning, particularly in Contextual Markov Decision Processes (CMDPs), where environments differ across episodes. Existing episodic intrinsic motivation methods for CMDPs primarily rely on count-based approac…

Cited by 0SourcePDFScholar
2025

Fewer May Be Better: Enhancing Offline Reinforcement Learning with Reduced Dataset

ICLR 2025poster

Research in offline reinforcement learning (RL) marks a paradigm shift in RL. However, a critical yet under-investigated aspect of offline RL is determining the subset of the offline dataset, which is used to improve algorithm performance while accelerating algorithm training. Moreover, the size of…

Cited by 0SourcePDFScholar
2025

Maximum Next-State Entropy for Efficient Reinforcement Learning

RA-L 2025

Entropy regularization is widely used to improve policy optimization and encourage exploration in reinforcement learning. By maximizing both the expected return and entropy, the agent aims to succeed at the task while acting as randomly as possible. However, current methods based on policy entropy e

Cited by 1SourceScholar
2024

Bayesian Design Principles for Offline-to-Online Reinforcement Learning

ICML 2024poster

Offline reinforcement learning (RL) is crucial for real-world applications where exploration can be costly or unsafe. However, offline learned policies are often suboptimal, and further online fine-tuning is required. In this paper, we tackle the fundamental dilemma of offline-to-online fine-tuning:…

2024

Learning Diverse Risk Preferences in Population-Based Self-Play

AAAI 2024technical

Among the remarkable successes of Reinforcement Learning (RL), self-play algorithms have played a crucial role in solving competitive games. However, current self-play RL methods commonly optimize the agent to maximize the expected win-rates against its current or historical copies, resulting in a l…

2024

No Prior Mask: Eliminate Redundant Action for Deep Reinforcement Learning

AAAI 2024technical

The large action space is one fundamental obstacle to deploying Reinforcement Learning methods in the real world. The numerous redundant actions will cause the agents to make repeated or invalid attempts, even leading to task failure. Although current algorithms conduct some initial explorations for…

2023

A Learning-Based Method for Computing Control Barrier Functions of Nonlinear Systems With Control Constraints

RA-L 2023

Verifying the safety of states and designing a safety controller are very important for safety critical systems (for example, robotic and automotive systems). Since the reachable set of a state is hard to calculate online, it is difficult to determine whether the current state will enter the unsafe

Cited by 2SourceScholar
2023

Flow to Control: Offline Reinforcement Learning with Lossless Primitive Discovery

AAAI 2023technical

Offline reinforcement learning (RL) enables the agent to effectively learn from logged data, which significantly extends the applicability of RL algorithms in real-world scenarios where exploration can be expensive or unsafe. Previous works have shown that extracting primitive skills from the recurr…

Cited by 18SourcePDFScholar
2023

Mean-Semivariance Policy Optimization via Risk-Averse Reinforcement Learning (Extended Abstract)

IJCAI 2023poster

Keeping risk under control is often more crucial than maximizing expected rewards in real-world decision-making situations, such as finance, robotics, autonomous driving, etc. The most natural choice of risk measures is variance, while it penalizes the upside volatility as much as the downside part.…

Cited by 0SourcePDFScholar
2023

The Provable Benefit of Unsupervised Data Sharing for Offline Reinforcement Learning

ICLR 2023poster

Self-supervised methods have become crucial for advancing deep learning by leveraging data itself to reduce the need for expensive annotations. However, the question of how to conduct self-supervised offline reinforcement learning (RL) in a principled way remains unclear. In this paper, we address t…

Cited by 17SourcePDFScholar
2022

Offline Reinforcement Learning with Value-based Episodic Memory

ICLR 2022poster

Offline reinforcement learning (RL) shows promise of applying RL to real-world problems by effectively utilizing previously collected data. Most existing offline RL algorithms use regularization or constraints to suppress extrapolation error for actions outside the dataset. In this paper, we adopt a…

Cited by 50SourcePDFScholar
2022

On the Role of Discount Factor in Offline Reinforcement Learning

ICML 2022spotlight

Offline reinforcement learning (RL) enables effective learning from previously collected data without exploration, which shows great promise in real-world applications when exploration is expensive or even infeasible. The discount factor, $\gamma$, plays a vital role in improving online RL sample ef…

Cited by 25SourcePDFScholar
2021

Average-Reward Reinforcement Learning with Trust Region Methods

IJCAI 2021poster

Most of reinforcement learning algorithms optimize the discounted criterion which is beneficial to accelerate the convergence and reduce the variance of estimates. Although the discounted criterion is appropriate for certain tasks such as financial related problems, many engineering problems treat f…

Cited by 22SourcePDFScholar
2021

Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement Learning

NeurIPS 2021spotlight

Learning from datasets without interaction with environments (Offline Learning) is an essential step to apply Reinforcement Learning (RL) algorithms in real-world scenarios. However, compared with the single-agent counterpart, offline multi-agent RL introduces more agents with the larger state and a…

2021

Celebrating Diversity in Shared Multi-Agent Reinforcement Learning

NeurIPS 2021poster

Recently, deep multi-agent reinforcement learning (MARL) has shown the promise to solve complex cooperative tasks. Its success is partly because of parameter sharing among agents. However, such sharing may lead agents to behave similarly and limit their coordination capacity. In this paper, we aim t…

Cited by 189SourcePDFScholar
2021

Optimal Dynamic Duct Static Pressure Method in a Multi-Zone Variable Air Volume System

RA-L 2021

Reducing the energy consumption of a variable air volume (VAV) system in heating, ventilation, and air-conditioning (HVAC) systems attracts many attentions. In this letter, a novel method, namely, optimal dynamic duct static pressure (ODSP) method is proposed to find the globally optimal solutions o

Cited by 5SourceScholar