← Search

Zhiwei Xu

30 accepted papers

2026

Accommodate Knowledge Conflicts in Retrieval-augmented LLMs: Towards Robust Response Generation in the Wild

AAAI 2026technical

The proliferation of large language models (LLMs) has significantly advanced intelligent systems. Unfortunately, LLMs often face knowledge conflicts between internal memory and retrieved external information, arising from misinformation, biases, or outdated knowledge. These conflicts undermine respo

Cited by 0SourcePDFScholar
2026

DRIVE: Distributional and Retrieval-Augmented Bidding with Value Evaluation

ICML 2026poster

Auto-bidding is a core component of real-time advertising systems, where decisions must optimize long-term performance under budget and cost constraints, while online exploration is prohibitively risky. Offline reinforcement learning and, more recently, Transformer-based sequence modeling have shown…

Cited by 1SourceScholar
2026

ExtendAttack: Attacking Servers of LRMs via Extending Reasoning

AAAI 2026technical

Large Reasoning Models (LRMs) have demonstrated promising performance in complex tasks. However, the resource-consuming reasoning processes may be exploited by attackers to maliciously occupy the resources of the servers, leading to a crash, like the DDoS attack in cyber. To this end, we propose a n

Cited by 0SourcePDFScholar
2026

From Traits to Roles: Consensus-Guided Composition of Orthogonal Experts for Cooperative MARL

IJCAI 2026

Parameter sharing is a central design choice in cooperative multi-agent reinforcement learning, yet it fundamentally conflicts with the need for role specialization in heterogeneous cooperative environments. Existing role-based methods typically learn monolithic role representations, which often suf

Cited by 0Scholar
2026

Graph of Verification: Structured Verification of LLM Reasoning with Directed Acyclic Graphs

AAAI 2026technical

Verifying the complex and multi-step reasoning of Large Language Models (LLMs) is a critical challenge, as holistic methods often overlook localized flaws. Step-by-step validation is a promising alternative, yet existing methods are often rigid. They struggle to adapt to diverse reasoning structur

Cited by 0SourcePDFScholar
2026

MultiGeo: Predicting Drug-Target Affinity via Adaptive Multi-Conformation Ensemble Learning

IJCAI 2026

Predicting drug–target affinity (DTA) is central to drug discovery, yet most deep learning models rely on a single static protein structure, neglecting the conformational heterogeneity that underlies many binding mechanisms. We propose MultiGeo, a DTA prediction framework that explicitly leverages m

Cited by 0Scholar
2026

Peak-Return Greedy Slicing: Subtrajectory Selection for Transformer-based Offline RL

ICLR 2026poster

Offline reinforcement learning enables policy learning solely from fixed datasets, without costly or risky environment interactions, making it highly valuable for real-world applications. While Transformer-based approaches have recently demonstrated strong sequence modeling capabilities, they typica…

Cited by 0SourceScholar
2026

RCVAFusion: 4D Radar and Camera Fusion With Virtual Points Association for 3D Object Detection

RA-L 2026

4D millimeter-wave radar plays a critical role in object detection for autonomous driving and robotics under all-weather and all-lighting conditions. Recently, the virtual-point-based approaches have attracted widespread attention due to their ability to address radar data sparsity by complementing

Cited by 0SourceScholar
2026

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

ICLR 2026poster

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 1.5: a carefully…

Cited by 0SourcecodeScholar
2025

Belief-Calibrated Multi-Agent Consensus Seeking for Complex NLP Tasks

NeurIPS 2025poster

A multi-agent system (MAS) enhances its capacity to solve complex natural language processing (NLP) tasks through collaboration among multiple agents, where consensus-seeking serves as a fundamental mechanism. However, existing consensus-seeking approaches typically rely on voting mechanisms to judg…

Cited by 0SourcecodeScholar
2025

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems

EMNLP 2025

The advancement of large language models (LLMs) has enabled the construction of multi-agent systems to solve complex tasks by dividing responsibilities among specialized agents, such as a planning agent for subgoal generation and a grounding agent for executing tool-use actions. Most existing method

2025

Efficient Communication in Multi-Agent Reinforcement Learning with Implicit Consensus Generation

AAAI 2025technical

A key challenge in multi-agent collaborative tasks is reducing uncertainty about teammates to enhance cooperative performance. Explicit communication methods can reduce uncertainty about teammates, but the associated high communication costs limit their practicality. Alternatively, implicit consensu…

Cited by 0SourcePDFScholar
2025

Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition

AAAI 2025technical

Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual…

2025

Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker Model

ICLR 2025poster

''Grokking'' is a phenomenon where a neural network first memorizes training data and generalizes poorly, but then suddenly transitions to near-perfect generalization after prolonged training. While intriguing, this delayed generalization phenomenon compromises predictability and efficiency. Ideally…

2025

Reidentify: Context-Aware Identity Generation for Contextual Multi-Agent Reinforcement Learning

ICML 2025poster

Generalizing multi-agent reinforcement learning (MARL) to accommodate variations in problem configurations remains a critical challenge in real-world applications, where even subtle differences in task setups can cause pre-trained policies to fail. To address this, we propose Context-Aware Identity…

Cited by 0SourcePDFScholar
2025

SLIM: Subtrajectory-Level Elimination for More Effective Reasoning

EMNLP 2025

In recent months, substantial progress has been made in complex reasoning of Large Language Models (LLMs), particularly through the application of test-time scaling. Notable examples include, though are not limited to, OpenAI’s o1/o3/o4 series and DeepSeek-R1. When responding to a query, these model

Cited by 0SourcePDFScholar
2024

Adaptive Parameter Sharing for Multi-Agent Reinforcement Learning

ICASSP 2024accepted

Parameter sharing, as an important technique in multi-agent systems, can effectively solve the scalability issue in large-scale agent problems. However, the effectiveness of parameter sharing largely depends on the environment setting. When agents have different identities or tasks, naive parameter…

Cited by 0SourceScholar
2024

Adversarial Purification with the Manifold Hypothesis

AAAI 2024technical

In this work, we formulate a novel framework for adversarial robustness using the manifold hypothesis. This framework provides sufficient conditions for defending against adversarial examples. We develop an adversarial purification method with this framework. Our method combines manifold learning wi…

2024

Benign Overfitting and Grokking in ReLU Networks for XOR Cluster Data

ICLR 2024poster

Neural networks trained by gradient descent (GD) have exhibited a number of surprising generalization behaviors. First, they can achieve a perfect fit to noisy training data and still generalize near-optimally, showing that overfitting can sometimes be benign. Second, they can undergo a period of cl…

Cited by 33SourcePDFScholar
2024

IMPUS: Image Morphing with Perceptually-Uniform Sampling Using Diffusion Models

ICLR 2024poster

We present a diffusion-based image morphing approach with perceptually-uniform sampling (IMPUS) that produces smooth, direct and realistic interpolations given an image pair. The embeddings of two images may lie on distinct conditioned distributions of a latent diffusion model, especially when they…

2024

Sequential Asynchronous Action Coordination in Multi-Agent Systems: A Stackelberg Decision Transformer Approach

ICML 2024poster

Asynchronous action coordination presents a pervasive challenge in Multi-Agent Systems (MAS), which can be represented as a Stackelberg game (SG). However, the scalability of existing Multi-Agent Reinforcement Learning (MARL) methods based on SG is severely restricted by network architectures or env…

Cited by 5SourcePDFScholar
2023

Consensus Learning for Cooperative Multi-Agent Reinforcement Learning

AAAI 2023technical

Almost all multi-agent reinforcement learning algorithms without communication follow the principle of centralized training with decentralized execution. During the centralized training, agents can be guided by the same signals, such as the global state. However, agents lack the shared signal and ch…

Cited by 17SourcePDFScholar
2023

Dual Self-Awareness Value Decomposition Framework without Individual Global Max for Cooperative MARL

NeurIPS 2023poster

Value decomposition methods have gained popularity in the field of cooperative multi-agent reinforcement learning. However, almost all existing methods follow the principle of Individual Global Max (IGM) or its variants, which limits their problem-solving capabilities. To address this, we propose a…

Cited by 4SourcePDFScholar
2023

HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination Mechanism

AAAI 2023technical

Recently, some challenging tasks in multi-agent systems have been solved by some hierarchical reinforcement learning methods. Inspired by the intra-level and inter-level coordination in the human nervous system, we propose a novel value decomposition framework HAVEN based on hierarchical reinforceme…

Cited by 33SourcePDFScholar
2023

Hierarchical Multi-Agent Reinforcement Learning with Intrinsic Reward Rectification

ICASSP 2023accepted

Hierarchical reinforcement learning (HRL) is a promising approach to solving long-term decision problems and complex tasks, as high-level policy can guide the training procedure of low-level policy with macro actions and intrinsic rewards. However, the amount that macro actions influence decision-ma…

Cited by 0SourceScholar
2023

Inducing Stackelberg Equilibrium through Spatio-Temporal Sequential Decision-Making in Multi-Agent Reinforcement Learning

IJCAI 2023poster

In multi-agent reinforcement learning (MARL), self-interested agents attempt to establish equilibrium and achieve coordination depending on game structure. However, existing MARL approaches are mostly bound by the simultaneous actions of all agents in the Markov game (MG) framework, and few works co…

Cited by 15SourcePDFScholar
2022

Mingling Foresight with Imagination: Model-Based Cooperative Multi-Agent Reinforcement Learning

NeurIPS 2022accept

Recently, model-based agents have achieved better performance than model-free ones using the same computational budget and training time in single-agent environments. However, due to the complexity of multi-agent systems, it is tough to learn the model of the environment. The significant compounding…

Cited by 13SourcePDFScholar