← Search

Xun Wang

49 accepted papers

2026

Benchmarking Empirical Privacy Protection for Adaptations of Large Language Models

ICLR 2026oral

Recent work has applied differential privacy (DP) to adapt large language models (LLMs) for sensitive applications, offering theoretical guarantees. However, its practical effectiveness remains unclear, partly due to LLM pretraining, where overlaps and interdependencies with adaptation data can unde…

Cited by 0SourceScholar
2026

Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation

CVPR 2026

Vision Transformers (ViTs) have recently achieved state-of-the-art performance in 2D human pose estimation due to their strong global modeling capability. However, existing ViT-based pose estimators are designed for static images and process each frame independently, thereby ignoring the temporal co

Cited by 0SourcecodeScholar
2026

Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training

ICML 2026poster

Generative Flow Networks (GFlowNets) excel at sampling diverse, high-reward objects. In many practical applications where active reward queries are infeasible, these models must be trained using static offline datasets. Prevailing training methods typically rely on a proxy model to provide reward fe…

Cited by 0SourceScholar
2026

CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image Retrieval

CVPR 2026

Interactive Text-to-Image Retrieval (I-TIR) aims to refine image retrieval results through natural language dialogues, which allows users to progressively supplement or correct their search intention across multiple rounds, enabling a more precise and user-aligned visual search experience.However, e

Cited by 0SourcecodeScholar
2026

Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies

ICML 2026poster

Online Multi-Agent Reinforcement Learning (MARL) is a prominent framework for efficient agent coordination. Crucially, enhancing policy expressiveness is pivotal for achieving superior performance. Diffusion-based generative models are well-positioned to meet this demand, having demonstrated remarka…

Cited by 0SourceScholar
2026

End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer

AAAI 2026technical

Existing multi-person video pose estimation methods typically adopt a two-stage pipeline: detecting individuals in each frame, followed by temporal modeling for single-person pose estimation. This design relies on heuristic operations such as tracking, RoI cropping, and non-maximum suppression, limi

Cited by 0SourcePDFScholar
2026

FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization

AAAI 2026technical

Garment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment

Cited by 0SourcePDFScholar
2025

Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

ACL 2025long

Language is not monolithic. While benchmarks, including those designed for multiple languages, are often used as proxies to evaluate the performance of Large Language Models (LLMs), they tend to overlook the nuances of within-language variation and thus fail to model the experience of speakers of no…

Cited by 0SourcePDFScholar
2025

Bridging Context Gaps: Leveraging Coreference Resolution for Long Contextual Understanding

ICLR 2025poster

Large language models (LLMs) have shown remarkable capabilities in natural language processing; however, they still face difficulties when tasked with understanding lengthy contexts and executing effective question answering. These challenges often arise due to the complexity and ambiguity present i…

2025

DP-MemArc: Differential Privacy Transfer Learning for Memory Efficient Language Models

AAAI 2025technical

Large language models have repeatedly shown outstanding performance across diverse applications. However, deploying these models can inadvertently risk user privacy. The significant memory demands during training pose a major challenge in terms of resource consumption. This substantial size places a…

Cited by 0SourcePDFScholar
2025

Developing a Reliable, Fast, General-Purpose Hallucination Detection and Mitigation Service

NAACL 2025industry

Hallucination, a phenomenon where large language models (LLMs) produce output that is factually incorrect or unrelated to the input, is a major challenge for LLM applications that require accuracy and dependability. In this paper, we introduce a reliable and high-speed production system aimed at det…

Cited by 0SourcePDFScholar
2025

Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval

AAAI 2025technical

Existing cross-modal retrieval methods typically rely on large-scale vision-language pair data. This makes it challenging to efficiently develop a cross-modal retrieval model for under-resourced languages of interest. Therefore, Cross-lingual Cross-modal Retrieval (CCR), which aims to align vision a…

2025

Efficient and Privacy-Preserving Soft Prompt Transfer for LLMs

ICML 2025poster

Prompting has become a dominant paradigm for adapting large language models (LLMs). While discrete (textual) prompts are widely used for their interpretability, soft (parameter) prompts have recently gained traction in APIs. This is because they can encode information from more training samples whil…

Cited by 0SourcePDFScholar
2025

K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning

NAACL 2025long

Strategic reasoning is a complex yet essential capability for intelligent agents. It requires Large Language Model (LLM) agents to adapt their strategies dynamically in multi-agent environments. Unlike static reasoning tasks, success in these contexts depends on anticipating other agents’ beliefs an…

Cited by 0SourcePDFScholar
2025

LLM-assisted Entropy-based Adaptive Distillation for Unsupervised Fine-grained Visual Representation Learning

ICCV 2025poster

Unsupervised Fine-grained Visual Represent Learning (FVRL) aims to learn discriminative features to distinguish subtle differences among visually similar categories without using labeled fine-grained data. Existing works, which typically learn representation from target data, often struggle to captu…

2025

OmniGen-AR: AutoRegressive Any-to-Image Generation

NeurIPS 2025poster

Autoregressive (AR) models have demonstrated strong potential in visual generation, offering competitive performance with simple architectures and optimization objectives. However, existing methods are typically limited to single-modality conditions, \eg, text or category labels, restricting their a…

Cited by 0SourceScholar
2025

Revisiting and Refining Lagunas' Beamforming for Acoustic Imaging

ICASSP 2025accepted

Lagunas et al proposed an adaptive beamforming method in 1986. However, this method never receives attention: until 2024, the citation of this paper is only 66. Actually, the Lagunas’ beamforming failed to identify sound sources if the required covariance matrix of array measurements is derived from…

Cited by 0SourceScholar
2025

Safe Planner: Empowering Safety Awareness in Large Pre-Trained Models for Robot Task Planning

AAAI 2025technical

Robot task planning is an important problem for autonomous robots in long-horizon challenging tasks. As large pre-trained models have demonstrated superior planning ability, recent research investigates utilizing large models to achieve autonomous planning for robots in diverse tasks. However, sinc…

Cited by 3SourcePDFScholar
2025

Tool-Planner: Task Planning with Clusters across Multiple Tools

ICLR 2025poster

Large language models (LLMs) have demonstrated exceptional reasoning capabilities, enabling them to solve various complex problems. Recently, this ability has been applied to the paradigm of tool learning. Tool learning involves providing examples of tool usage and their corresponding functions, all…

2025

Towards Ship License Plate Recognition in the Wild: A Large Benchmark and Strong Baseline

AAAI 2025technical

The paper targets the challenging task of Ship License Plate (SLP) recognition. Existing methods for SLP recognition are hampered by the scarcity of large and publicly available datasets, leading to evaluations on small and non-representative datasets. To alleviate it, we have built a large dataset,…

2024

Continual Vision-Language Retrieval via Dynamic Knowledge Rectification

AAAI 2024technical

The recent large-scale pre-trained models like CLIP have aroused great concern in vision-language tasks. However, when required to match image-text data collected in a streaming manner, namely Continual Vision-Language Retrieval (CVRL), their performances are still limited due to the catastrophic fo…

2024

In-context Autoencoder for Context Compression in a Large Language Model

ICLR 2024poster

We propose the In-context Autoencoder (ICAE), leveraging the power of a large language model (LLM) to compress a long context into short compact memory slots that can be directly conditioned on by the LLM for various purposes. ICAE is first pretrained using both autoencoding and language modeling ob…

2024

Learning Continual Compatible Representation for Re-indexing Free Lifelong Person Re-identification

CVPR 2024poster

Lifelong Person Re-identification (L-ReID) aims to learn from sequentially collected data to match a person across different scenes. Once an L-ReID model is updated using new data all historical images in the gallery are required to be re-calculated to obtain new features for testing known as "re-in…

2024

NoteChat: A Dataset of Synthetic Patient-Physician Conversations Conditioned on Clinical Notes

ACL 2024findings

We introduce NoteChat, a novel cooperative multi-agent framework leveraging Large Language Models (LLMs) to generate patient-physician dialogues. NoteChat embodies the principle that an ensemble of role-specific LLMs, through structured role-play and strategic prompting, can perform their assigned r…

2024

ODD: A Benchmark Dataset for the Natural Language Processing Based Opioid Related Aberrant Behavior Detection

NAACL 2024long

Opioid related aberrant behaviors (ORABs) present novel risk factors for opioid overdose. This paper introduces a novel biomedical natural language processing benchmark dataset named ODD, for ORAB Detection Dataset. ODD is an expert-annotated dataset designed to identify ORABs from patients’ EHR not…

2024

Robust Visual Imitation Learning with Inverse Dynamics Representations

AAAI 2024technical

Imitation learning (IL) has achieved considerable success in solving complex sequential decision-making problems. However, current IL methods mainly assume that the environment for learning policies is the same as the environment for collecting expert datasets. Therefore, these methods may fail to w…

Cited by 2SourcePDFScholar
2024

SCALE: Synergized Collaboration of Asymmetric Language Translation Engines

ACL 2024findings

In this paper, we introduce SCALE, a collaborative framework that connects a compact Specialized Translation Model (STM) and a general-purpose Large Language Model (LLM) as one unified translation engine. By introducing translation from STM into the triplet in-context demonstrations, SCALE unlocks r…

2024

SecCoder: Towards Generalizable and Robust Secure Code Generation

EMNLP 2024main

After large models (LMs) have gained widespread acceptance in code-related tasks, their superior generative capacity has greatly promoted the application of the code LM. Nevertheless, the security of the generated code has raised attention to its potential damage. Existing secure code generation met…

Cited by 0SourcePDFScholar
2024

Simultaneous Optimization of Bid Shading and Internal Auction for Demand-Side Platforms

AAAI 2024technical

Online advertising has been one of the most important sources for industry's growth, where the demand-side platforms (DSP) play an important role via bidding to the ad exchanges on behalf of their advertiser clients. Since more and more ad exchanges have shifted from second to first price auctions,…

Cited by 5SourcePDFScholar
2024

Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement

EMNLP 2024main

Large language model agents have exhibited exceptional performance across a range of complex interactive tasks. Recent approaches have utilized tuning with expert trajectories to enhance agent performance, yet they primarily concentrate on outcome rewards, which may lead to errors or suboptimal acti…

2024

xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token

NeurIPS 2024poster

This paper introduces xRAG, an innovative context compression method tailored for retrieval-augmented generation. xRAG reinterprets document embeddings in dense retrieval--traditionally used solely for retrieval--as features from the retrieval modality. By employing a modality fusion methodology, xR…

2023

Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video Retrieval

ICCV 2023poster

Almost all previous text-to-video retrieval works assume that videos are pre-trimmed with short durations. However, in practice, videos are generally untrimmed containing much background content. In this work, we investigate the more practical but challenging Partially Relevant Video Retrieval (PRVR…

Cited by 19PDFcodeScholar
2023

Extensible Prompts for Language Models on Zero-shot Language Style Customization

NeurIPS 2023poster

We propose eXtensible Prompt (X-Prompt) for prompting a large language model (LLM) beyond natural language (NL). X-Prompt instructs an LLM with not only NL but also an extensible vocabulary of imaginary words. Registering new imaginary words allows us to instruct the LLM to comprehend concepts that…

Cited by 3SourcePDFScholar
2023

Hierarchical Contrast for Unsupervised Skeleton-Based Action Representation Learning

AAAI 2023technical

This paper targets unsupervised skeleton-based action representation learning and proposes a new Hierarchical Contrast (HiCo) framework. Different from the existing contrastive-based solutions that typically represent an input skeleton sequence into instance-level features and perform contrast holis…

2023

Hypothesis Test for Leakage Detection in Water Pipelines with High-Dimensional Sensor Signals

ICASSP 2023accepted

We design a statistical hypothesis test for performing leak detection in water pipeline channels. By applying an appropriate model for signal propagation, we show that the detection problem becomes one of distinguishing signal from noise, with the noise being described by a multivariate Gaussian dis…

Cited by 0SourceScholar
2023

Sequence-Based Device-Free Gesture Recognition Framework for Multi-Channel Acoustic Signals

ICASSP 2023accepted

Device-free gesture recognition schemes based on acoustic sensing signals are promising solutions for next-generation human-computer interaction systems. However, existing gesture recognition frameworks reuse visual neural networks to perform feature extraction. These approaches ignore the time sequ…

Cited by 0SourceScholar
2023

Smart Word Suggestions for Writing Assistance

ACL 2023findings

Enhancing word usage is a desired feature for writing assistance. To further advance research in this area, this paper introduces “Smart Word Suggestions” (SWS) task and benchmark. Unlike other works, SWS emphasizes end-to-end evaluation and presents a more realistic writing assistance scenario. Thi…

2023

Two Directions for Clinical Data Generation with Large Language Models: Data-to-Label and Label-to-Data

EMNLP 2023long findings

Large language models (LLMs) can generate natural language texts for various domains and tasks, but their potential for clinical text mining, a domain with scarce, sensitive, and imbalanced medical data, is under-explored. We investigate whether LLMs can augment clinical data for detecting Alzheimer…

Cited by 0SourceScholar
2021

Deep Dual Consecutive Network for Human Pose Estimation

CVPR 2021poster

Multi-frame human pose estimation in complicated situations is challenging. Although state-of-the-art human joints detectors have demonstrated remarkable results for static images, their performances come short when we apply these models to video sequences. Prevalent shortcomings include the failure…

Cited by 166PDFcodeScholar
2021

Spatial Ensemble: a Novel Model Smoothing Mechanism for Student-Teacher Framework

NeurIPS 2021poster

Model smoothing is of central importance for obtaining a reliable teacher model in the student-teacher framework, where the teacher generates surrogate supervision signals to train the student. A popular model smoothing method is the Temporal Moving Average (TMA), which continuously averages the tea…

2019

Dual Encoding for Zero-Example Video Retrieval

CVPR 2019poster

This paper attacks the challenging problem of zero-example video retrieval. In such a retrieval paradigm, an end user searches for unlabeled videos by ad-hoc queries described in natural language text with no visual example provided. Given videos as sequences of frames and queries as sequences of wo…

Cited by 324PDFcodeScholar
2019

Multi-Similarity Loss With General Pair Weighting for Deep Metric Learning

CVPR 2019poster

A family of loss functions built on pair-based computation have been proposed in the literature which provide a myriad of solutions for deep metric learning. In this pa-per, we provide a general weighting framework for under-standing recent pair-based loss functions. Our contributions are t…

Cited by 1000PDFcodeScholar