← Search

Xiao Wu

25 accepted papers

2026

I2CD: An Invertible Causal Framework for Compositional Zero-Shot Learning via Disentangle-Compose-Disentangle

AAAI 2026technical

Compositional Zero-Shot Learning (CZSL) addresses the challenge of recognizing unseen attribute-object compositions in images, representing a fundamental challenge in artificial intelligence. Current approaches, which primarily focus on semantic alignment or distribution independence of primitives,

Cited by 0SourcePDFScholar
2026

InfoCom: Kilobyte-Scale Communication-Efficient Collaborative Perception with Information Bottleneck

AAAI 2026technical

Precise environmental perception is critical for the reliability of autonomous driving systems. While collaborative perception mitigates the limitations of single-agent perception through information sharing, it encounters a fundamental communication-performance trade-off. Existing communication-eff

Cited by 0SourcePDFScholar
2026

Monocular Vehicle Pose and Shape Reconstruction via Dynamic Context Adaptation and Progressive Geometry Refinement

AAAI 2026technical

Accurate reconstruction of 3D vehicle pose and shape from monocular images is challenging, particularly for distant objects in autonomous driving. Existing methods often suffer from geometric ambiguity in depth estimation and structural hollowness in shape recovery, primarily due to inadequate multi

Cited by 0SourcePDFScholar
2026

Optical Flow Matching: Reframing Optical Flow as Continuous Transport Dynamics

CVPR 2026

Modern optical flow estimation, though empowered by recent deep neural architectures, remains rooted in the discrete correspondence paradigm inherited from classical vision. Most networks infer frame-to-frame displacements, capturing where pixels move but not how motion evolves continuously through

Cited by 0SourcecodeScholar
2025

3R: Enhancing Sentence Representation Learning via Redundant Representation Reduction

EMNLP 2025

Sentence representation learning (SRL) aims to learn sentence embeddings that conform to the semantic information of sentences. In recent years, fine-tuning methods based on pre-trained models and contrastive learning frameworks have significantly advanced the quality of sentence representations. Ho

2025

A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-making

EMNLP 2025

Medical decision-making often involves integrating knowledge from multiple clinical specialties, typically achieved through multidisciplinary teams. Inspired by this collaborative process, recent work has leveraged large language models (LLMs) in multi-agent collaboration frameworks to emulate exper

2025

CoPEFT: Fast Adaptation Framework for Multi-Agent Collaborative Perception with Parameter-Efficient Fine-Tuning

AAAI 2025technical

Multi-agent collaborative perception is expected to significantly improve perception performance by overcoming the limitations of single-agent perception through exchanging complementary information. However, training a robust collaborative perception model requires collecting sufficient training da…

2025

Exploring Temporal Constraints for Unsupervised Iris Motion Tracking in AS-OCT Videos

ICASSP 2025accepted

Iris motion tracking is critical for discriminating the iris stiffness and developmental stage of primary angle-closure disease (PACD). Anterior segment optical coherence tomography (AS-OCT) video is a highly efficient approach to observe the morphological determinant in iris motion. However, the ir…

Cited by 0SourceScholar
2025

Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-Learning

AAAI 2025technical

Dense video captioning aims to detect and describe all events in untrimmed videos. This paper presents a dense video captioning network called Multi-Concept Cyclic Learning (MCCL), which aims to: (1) detect multiple concepts at the frame level and leverage these concepts to provide temporal event cu…

Cited by 0SourcePDFScholar
2025

POPoS: Improving Efficient and Robust Facial Landmark Detection with Parallel Optimal Position Search

AAAI 2025technical

Achieving a balance between accuracy and efficiency is a critical challenge in facial landmark detection (FLD). This paper introduces Parallel Optimal Position Search (POPoS), a high-precision encoding-decoding framework designed to address the limitations of traditional FLD methods. POPoS employs t…

2025

Pragmatic Heterogeneous Collaborative Perception via Generative Communication Mechanism

NeurIPS 2025poster

Multi-agent collaboration enhances the perception capabilities of individual agents through information sharing. However, in real-world applications, differences in sensors and models across heterogeneous agents inevitably lead to domain gaps during collaboration. Existing approaches based on adapta…

Cited by 0SourcecodeScholar
2024

Content-Adaptive Non-Local Convolution for Remote Sensing Pansharpening

CVPR 2024poster

Currently machine learning-based methods for remote sensing pansharpening have progressed rapidly. However existing pansharpening methods often do not fully exploit differentiating regional information in non-local spaces thereby limiting the effectiveness of the methods and resulting in redundant l…

2024

Exploring the Low-Pass Filtering Behavior in Image Super-Resolution

ICML 2024poster

Deep neural networks for image super-resolution (ISR) have shown significant advantages over traditional approaches like the interpolation. However, they are often criticized as 'black boxes' compared to traditional approaches with solid mathematical foundations. In this paper, we attempt to interpr…

2024

PostureHMR: Posture Transformation for 3D Human Mesh Recovery

CVPR 2024poster

Human Mesh Recovery (HMR) aims to estimate the 3D human body from 2D images which is a challenging task due to inherent ambiguities in translating 2D observations to 3D space. A novel approach called PostureHMR is proposed to leverage a multi-step diffusion-style process which converts this task int…

Cited by 3SourcePDFScholar
2024

SSDiff: Spatial-spectral Integrated Diffusion Model for Remote Sensing Pansharpening

NeurIPS 2024poster

Pansharpening is a significant image fusion technique that merges the spatial content and spectral characteristics of remote sensing images to generate high-resolution multispectral images. Recently, denoising diffusion probabilistic models have been gradually applied to visual tasks, enhancing cont…

2023

Bidirectional Dilation Transformer for Multispectral and Hyperspectral Image Fusion

IJCAI 2023poster

Transformer-based methods have proven to be effective in achieving long-distance modeling, capturing the spatial and spectral information, and exhibiting strong inductive bias in various computer vision tasks. Generally, the Transformer model includes two common modes of multi-head self-attention (M…

Cited by 19SourcePDFScholar
2022

A Decoder-free Transformer-like Architecture for High-efficiency Single Image Deraining

IJCAI 2022poster

Despite the success of vision Transformers for the image deraining task, they are limited by computation-heavy and slow runtime. In this work, we investigate Transformer decoder is not necessary and has huge computational costs. Therefore, we revisit the standard vision Transformer as well as its su…

2022

Learning Graph-based Residual Aggregation Network for Group Activity Recognition

IJCAI 2022poster

Group activity recognition aims to understand the overall behavior performed by a group of people. Recently, some graph-based methods have made progress by learning the relation graphs among multiple persons. However, the differences between an individual and others play an important role in identif…

Cited by 9SourcePDFScholar
2022

Rethinking Spatial Invariance of Convolutional Networks for Object Counting

CVPR 2022poster

Previous work generally believes that improving the spatial invariance of convolutional networks is the key to object counting. However, after verifying several mainstream counting networks, we surprisingly found too strict pixel-level spatial invariance would cause overfit noise in the density map…

Cited by 124PDFcodeScholar
2022

SpanConv: A New Convolution via Spanning Kernel Space for Lightweight Pansharpening

IJCAI 2022poster

Standard convolution operations can effectively perform feature extraction and representation but result in high computational cost, largely due to the generation of the original convolution kernel corresponding to the channel dimension of the feature map, which will cause unnecessary redundancy. In…

2019

Generative Adversarial Networks Based Error Concealment for Low Resolution Video

ICASSP 2019accepted

In this paper, a novel deep generative model-based approach for video error concealment is proposed. Our method is comprised of completion network and two critics. The frame completion network is trained to fool the both the local and global critics, which requires completion network to conceal fram…

Cited by 0SourceScholar
2019

Learning Spatial Awareness to Improve Crowd Counting

ICCV 2019oral

The aim of crowd counting is to estimate the number of people in images by leveraging the annotation of center positions for pedestrians' heads. Promising progresses have been made with the prevalence of deep Convolutional Neural Networks. Existing methods widely employ the Euclidean distance (i.e.,…

Cited by 162PDFcodeScholar
2017

Memory-Augmented Attribute Manipulation Networks for Interactive Fashion Search

CVPR 2017poster

We introduce a new fashion search protocol where attribute manipulation is allowed within the interaction between users and search engines, e.g. manipulating the color attribute of the clothing from red to blue. It is particularly useful for image-based search when the query image cannot perfectly m…

Cited by 173PDFScholar
2017

Video2Shop: Exact Matching Clothes in Videos to Online Shopping Images

CVPR 2017poster

In recent years, both online retail and video hosting service have been exponentially grown. In this paper, a novel deep neural network, called AsymNet, is proposed to explore a new cross-domain task, Video2Shop, targeting for matching clothes appeared in videos to the exactly same items in online s…

Cited by 108PDFcodeScholar