← Search

Yuanzhao Zhai

12 accepted papers

2026

CoPE: A Framework for Optimizing Coordination between Planning and Execution in LLM-based Agents

ICML 2026poster

Fine-tuning Large Language Models (LLMs) as autonomous agents on domain-specific data has emerged as a promising paradigm for tackling interactive, real-world tasks. However, existing studies have overlooked the critical coordination between long-term planning and multi-step execution in optimizing …

Cited by 0SourceScholar
2025

COPR: Continual Human Preference Learning via Optimal Policy Regularization

ACL 2025finding

Reinforcement Learning from Human Feedback (RLHF) is effective for aligning Large Language Models (LLMs) with human preferences. However, RLHF’s complex process limits its ability to continually learn human feedback, making it impractical for real-world applications where the deployed model continuo…

Cited by 0SourcePDFScholar
2025

Correcting Large Language Model Behavior via Influence Function

AAAI 2025technical

Recent advancements in AI alignment techniques have significantly improved the alignment of large language models (LLMs) with static human preferences. However, the dynamic nature of human preferences can render some prior training data outdated or even erroneous, ultimately causing LLMs to deviate…

Cited by 0SourcePDFScholar
2025

Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

AAAI 2025technical

Agents significantly enhance the capabilities of standalone Large Language Models (LLMs) by perceiving environments, making decisions, and executing actions. However, LLM agents still face challenges in tasks that require multiple decision-making steps. Estimating the value of actions in specific ta…

Cited by 8SourcePDFScholar
2024

C3F: Constant Collaboration and Communication Framework for Graph-Representation Dynamic Multi-Robotic Systems

RA-L 2024

Deep reinforcement learning (DRL) methods have been widely applied in distributed multi-robotic systems and successfully realized autonomous learning in many fields. In these fields, robots need to communicate and collaborate with other robots in real time, and reach agreed cognition for task assign

Cited by 0SourceScholar
2024

Iterative Regularized Policy Optimization with Imperfect Demonstrations

ICML 2024poster

Imitation learning heavily relies on the quality of provided demonstrations. In scenarios where demonstrations are imperfect and rare, a prevalent approach for refining policies is through online fine-tuning with reinforcement learning, in which a Kullback–Leibler (KL) regularization is often employ…

2024

Optimistic Model Rollouts for Pessimistic Offline Policy Optimization

AAAI 2024technical

Model-based offline reinforcement learning (RL) has made remarkable progress, offering a promising avenue for improving generalization with synthetic model rollouts. Existing works primarily focus on incorporating pessimism for policy optimization, usually via constructing a Pessimistic Markov Decis…

Cited by 1SourcePDFScholar
2023

Diversifying Message Aggregation in Multi-Agent Communication Via Normalized Tensor Nuclear Norm Regularization

ICASSP 2023accepted

The use of graph attention networks (GAT) in communication-enhanced multi-agent reinforcement learning (Comm-MARL) has become prevalent. While successful, GAT can lead to homogeneity in the strategies of message aggregation, which can severely limit multi-agent coordination. To address this challeng…

Cited by 0SourceScholar
2023

Progressive Diversifying Policy for Multi-Agent Reinforcement Learning

ICASSP 2023accepted

Multi-Agent Reinforcement Learning (MARL) has recently achieved promising performance in many collaborative decision making tasks. However, one of the main bottleneck challenges for MARL is the sparsity of the team reward, which can lead to the homogenization of agents’ behaviors. To address these i…

Cited by 0SourceScholar
2022

CRMRL: Collaborative Relationship Meta Reinforcement Learning for Effectively Adapting to Type Changes in Multi-Robotic System

RA-L 2022

Multi-agent reinforcement learning methods have been widely used for multi-robotic systems, and meta-learning methods are also applied to help robots reuse prior experiences to guide new tasks learning. But in some multi-robotic tasks, the robot types cannot be determined in advance or may dynamical

Cited by 7SourceScholar
2021

Decentralized Multi-Robot Collision Avoidance in Complex Scenarios With Selective Communication

RA-L 2021

Deep reinforcement learning has been demonstrated to be an effective solution to the multi-robot collision avoidance problem. However, with existing methods, robots typically generate actions only based on local observations, sometimes augmented with global communication. Their performance deteriora

Cited by 28SourceScholar