← Search

Xiangfeng Wang

19 accepted papers

2026

CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement

ICML 2026poster

While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conversation-level collaborative studies lack grounded interaction and behavioral execution, motivating the need for cooperative game environments that en…

Cited by 0SourceScholar
2026

Negotiated Reasoning: On Provably Addressing Relative Over-Generalization

ICLR 2026poster

We focus on the relative over-generalization (RO) issue in fully cooperative multi-agent reinforcement learning (MARL). Existing methods show that endowing agents with reasoning can help mitigate RO empirically, but there is little theoretical insight. We first prove that RO is avoided when agents s…

Cited by 0SourceScholar
2026

Who Matters Matters: Agent-Specific Conservative Offline MARL

ICLR 2026poster

Offline Multi-Agent Reinforcement Learning (MARL) enables policy learning from static datasets in multi-agent systems, eliminating the need for risky or costly environment interactions during training. A central challenge in offline MARL lies in achieving effective collaboration among heterogeneous…

Cited by 0SourceScholar
2025

LOPT: Learning Optimal Pigovian Tax in Sequential Social Dilemmas

NeurIPS 2025poster

Multi-agent reinforcement learning (MARL) has emerged as a powerful framework for modeling autonomous agents that independently optimize their individual objectives. However, in mixed-motive MARL environments, rational self-interested behaviors often lead to collectively suboptimal outcomes situatio…

Cited by 0SourceScholar
2025

Multi-Agent Credit Assignment with Pretrained Language Models

AISTATS 2025poster

The difficulty of appropriately assigning credit is particularly heightened in cooperative MARL with sparse reward, due to the concurrent time and structural scales involved. Automatic subgoal generation (ASG) has recently emerged as a viable MARL approach inspired by utilizing subgoals in intrinsic…

Cited by 0SourceScholar
2025

Reward Translation via Reward Machine in Semi-Alignable MDPs

ICML 2025poster

Addressing reward design complexities in deep reinforcement learning is facilitated by knowledge transfer across different domains. To this end, we define \textit{reward translation} to describe the cross-domain reward transfer problem. However, current methods struggle with non-pairable and non-tim…

Cited by 0SourcePDFScholar
2025

Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents

NeurIPS 2025poster

Large Language Models (LLMs) are increasingly deployed as autonomous agents in multi-agent systems, and promising coordination has been demonstrated in handling complex tasks under predefined roles and scripted workflows. However, significant challenges remain in open-ended environments, where agent…

Cited by 0SourceScholar
2025

SkyRover: A Modular Simulator for Cross-Domain Pathfinding

IJCAI 2025

Unmanned Aerial Vehicles (UAVs) and Automated Guided Vehicles (AGVs) increasingly collaborate in logistics, surveillance, inspection tasks and etc. However, existing simulators often focus on a single domain, limiting cross-domain study. This paper presents the SkyRover, a modular simulator for UAV-

2024

In-Context Former: Lightning-fast Compressing Context for Large Language Model

EMNLP 2024finding

With the rising popularity of Transformer-based large language models (LLMs), reducing their high inference costs has become a significant research focus. One effective approach to mitigate these costs is compressing the long input contexts. Existing methods typically leverage the self-attention mec…

2024

Long-tailed Diffusion Models with Oriented Calibration

ICLR 2024poster

Diffusion models are acclaimed for generating high-quality and diverse images. However, their performance notably degrades when trained on data with a long-tailed distribution. For long tail diffusion model generation, current works focus on the calibration and enhancement of the tail generation wit…

2022

Dealing with Non-Stationarity in MARL via Trust-Region Decomposition

ICLR 2022poster

Non-stationarity is one thorny issue in cooperative multi-agent reinforcement learning (MARL). One of the reasons is the policy changes of agents during the learning process. Some existing works have discussed various consequences caused by non-stationarity with several kinds of measurement indicato…

Cited by 19SourcePDFScholar
2022

Multi-Agent Path Finding with Prioritized Communication Learning

ICRA 2022poster

Multi-agent pathfinding (MAPF) has been widely used to solve large-scale real-world problems, e.g., automation warehouses. The learning-based, fully decentralized framework has been introduced to alleviate real-time problems and simultaneously pursue optimal planning policy. However, existing method…

Cited by 57SourcecodeScholar
2022

VMAgent: A Practical Virtual Machine Scheduling Platform

IJCAI 2022poster

Virtual machine (VM) scheduling is one of the critical tasks in cloud computing. Many works have attempted to incorporate machine learning, especially reinforcement learning, to empower VM scheduling procedures. Although improved results are shown in several demo simulators, the performances in real…

2020

Iteratively-Refined Interactive 3D Medical Image Segmentation With Multi-Agent Reinforcement Learning

CVPR 2020poster

Existing automatic 3D image segmentation methods usually fail to meet the clinic use. Many studies have explored an interactive strategy to improve the image segmentation performance by iteratively incorporating user hints. However, the dynamic process for successive interactions is largely ignored.…

Cited by 130PDFScholar
2019

A Fast Proximal Point Method for Computing Exact Wasserstein Distance

UAI 2019poster

Wasserstein distance plays increasingly important roles in machine learning, stochastic programming and image processing. Major efforts have been under way to address its high computational complexity, some leading to approximate or regularized variations such as Sinkhorn distance. However, as we wi…

Cited by 224SourcePDFScholar
2016

Asynchronous distributed alternating direction method of multipliers: Algorithm and convergence analysis

ICASSP 2016accepted

Alternating direction method of multipliers (ADMM) has been recognized as an efficient approach for solving many large-scale learning problems over a computer cluster. However, traditional synchronized computation does not scale well with the problem size, as the speed of the algorithm is limited by…

Cited by 0SourceScholar
2016

Nonnegative matrix factorization using ADMM: Algorithm and convergence analysis

ICASSP 2016accepted

The nonnegative matrix factorization (NMF) has been a popular model for a wide range of signal processing and machine learning problems. It is usually formulated as a nonconvex cost minimization problem. This work settles the convergence issue of a popular algorithm based on the alternating directio…

Cited by 0SourceScholar