← Search

Zhiqiang Pu

13 accepted papers

2026

Unreal-MAP: Unreal-Engine-Based General Platform for Multi-agent Reinforcement Learning

AAAI 2026technical

In this paper, we propose Unreal Multi-Agent Playground (Unreal-MAP), an MARL general platform based on the Unreal-Engine (UE). Unreal-MAP allows users to freely create multi-agent tasks using the vast visual and physical resources available in the UE community, and deploy state-of-the-art (SOTA) MA

Cited by 0SourcePDFScholar
2025

CLGA: A Collaborative LLM Framework for Dynamic Goal Assignment in Multi-Robot Systems

IROS 2025

Goal assignment is a critical challenge in multi-robot systems. The emergence of large language models (LLMs) has enabled the use of natural language commands for tackling goal assignment problems. However, applying LLMs directly to these tasks presents two limitations: 1) limited accuracy and 2) ex

Cited by 0SourceScholar
2025

CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning

EMNLP 2025

In parameter-efficient fine-tuning, mixture-of-experts (MoE), which involves specializing functionalities into different experts and sparsely activating them appropriately, has been widely adopted as a promising approach to trade-off between model capacity and computation overhead. However, current

Cited by 0SourcePDFScholar
2025

Stochastic Trajectory Prediction Under Unstructured Constraints

ICRA 2025

Trajectory prediction facilitates effective planning and decision-making, while constrained trajectory prediction integrates regulation into prediction. Recent advances in constrained trajectory prediction focus on structured constraints by constructing optimization objectives. However, handling uns

Cited by 2SourceScholar
2025

Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement Learning

AAAI 2025technical

Guiding the policy of multi-agent reinforcement learning to align with human common sense is a difficult problem, largely due to the complexity of modeling common sense as a reward, especially in complex and long-horizon multi-agent tasks. Recent works have shown the effectiveness of reward shaping,…

2024

Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning

NeurIPS 2024poster

Reinforcement learning (RL) has emerged as a pivotal technique for fine-tuning large language models (LLMs) on specific tasks. However, prevailing RL fine-tuning methods predominantly rely on PPO and its variants. Though these algorithms are effective in general RL settings, they often exhibit subop…

2023

Deconfounded Opponent Intention Inference for Football Multi-Player Policy Learning

IROS 2023poster

Due to the high complexity of a football match, the opponents' strategies are variable and unknown. Thus predicting the opponents' future intentions accurately based on current situation is crucial for football players' decision-making. To better anticipate the opponents and learn more effective str…

Cited by 2SourceScholar
2023

Lazy Agents: A New Perspective on Solving Sparse Reward Problem in Multi-agent Reinforcement Learning

ICML 2023poster

Sparse reward remains a valuable and challenging problem in multi-agent reinforcement learning (MARL). This paper addresses this issue from a new perspective, i.e., lazy agents. We empirically illustrate how lazy agents damage learning from both exploration and exploitation. Then, we propose a novel…

2022

Concentration Network for Reinforcement Learning of Large-Scale Multi-Agent Systems

AAAI 2022technical

When dealing with a series of imminent issues, humans can naturally concentrate on a subset of these concerning issues by prioritizing them according to their contributions to motivational indices, e.g., the probability of winning a game. This idea of concentration offers insights into reinforcement…

2022

Multi-Target Encirclement with Collision Avoidance via Deep Reinforcement Learning using Relational Graphs

ICRA 2022poster

In this paper, we propose a novel decentralized method based on deep reinforcement learning using robot-level and target-level relational graphs, to solve the problem of multi-target encirclement with collision avoidance (MECA). Specifically, the robot-level relational graphs, composed of three hete…

Cited by 13SourceScholar
2022

Multi-UAV Cooperative Short-Range Combat via Attention-Based Reinforcement Learning using Individual Reward Shaping

IROS 2022poster

In this paper, we propose a novel distributed method based on attention-based deep reinforcement learning using individual reward shaping, for multiple unmanned aerial vehicles (UAVs) cooperative short-range combat mission. Specifically, a two-level attention distributed policy, composed of observat…

Cited by 13SourceScholar
2021

Multi-agent Collaborative Learning with Relational Graph Reasoning in Adversarial Environments

IROS 2021poster

This paper proposes a collaborative policy framework via relational graph reasoning for multi-agent systems to accomplish adversarial tasks. A relational graph reasoning module consisting of an agent graph reasoning module and an opponent graph module, is designed to enable each agent to learn mixtu…

Cited by 8SourceScholar
2021

Multi-target Coverage with Connectivity Maintenance using Knowledge-incorporated Policy Framework

ICRA 2021poster

This paper considers a multi-target coverage problem where a robot team aims to efficiently cover multi-targets while maintaining connectivity in a distributed manner. A novel knowledge-incorporated policy framework is proposed to derive a distributed, efficient, and connectivity guaranteed coverage…

Cited by 11SourceScholar