← Search

Xiang Wu

15 accepted papers

2026

Scalable Topology-Preserving Graph Coarsening: Concepts and Algorithms

ICML 2026poster

Graph coarsening reduces the size of a graph while preserving certain properties. Most existing methods preserve either spectral or spatial characteristics. Recent research has shown that preserving topological features helps maintain the predictive performance of graph neural networks (GNNs) traine…

Cited by 0SourceScholar
2024

FusedNet: End-to-End Mobile Robot Relocalization in Dynamic Large-Scale Scene

RA-L 2024

To improve robot relocalization accuracy in both static and dynamic environments, we introduce a novel network, FusedNet, which incorporates a cross-attention to fuse global and local image features for end-to-end relocalization. This approach relies solely on a monocular camera sensor that is fixed

Cited by 5SourceScholar
2024

GRID: Scene-Graph-based Instruction-driven Robotic Task Planning

IROS 2024

Recent works have shown that Large Language Models (LLMs) can facilitate the grounding of instructions for robotic task planning. Despite this progress, most existing works have primarily focused on utilizing raw images to aid LLMs in understanding environmental information. However, this approach n

Cited by 45SourcecodeScholar
2021

CM-NAS: Cross-Modality Neural Architecture Search for Visible-Infrared Person Re-Identification

ICCV 2021poster

Visible-Infrared person re-identification (VI-ReID) aims to match cross-modality pedestrian images, breaking through the limitation of single-modality person ReID in dark environment. In order to mitigate the impact of large modality discrepancy, existing works manually design various two-stream arc…

Cited by 159PDFcodeScholar
2020

HAMBox: Delving Into Mining High-Quality Anchors on Face Detection

CVPR 2020poster

Current face detectors utilize anchors to frame a multi-task learning problem which combines classification and bounding box regression. Effective anchor design and anchor matching strategy enable face detectors to localize faces under large pose and scale variations. However, we observe that, more…

Cited by 50PDFScholar
2020

Hierarchical Face Aging through Disentangled Latent Characteristics

ECCV 2020poster

Current age datasets lie in a long-tailed distribution, which brings difficulties to describe the aging mechanism for the imbalance ages. To alleviate it, we design a novel facial age prior to guide the aging mechanism modeling. To explore the age effects on facial images, we propose a Disentangled…

Cited by 25SourcePDFScholar
2020

Hierarchical Sequence Representation with Graph Network

ICASSP 2020accepted

Video classification problem is a challenging task in computer vision. The performance of this task is highly relied on the scale of training data and the effectiveness of video embedding via a robust embedding network. Unsupervised solutions such as feature average pooling technique, as a simple la…

Cited by 0SourceScholar
2020

TF-NAS: Rethinking Three Search Freedoms of Latency-Constrained Differentiable Neural Architecture Search

ECCV 2020poster

With the flourish of differentiable neural architecture search (NAS), automatically searching latency-constrained architectures gives a new perspective to reduce human labor and expertise. However, the searched architectures are usually suboptimal in accuracy and may have large jitters around the ta…

2019

Breaking the Glass Ceiling for Embedding-Based Classifiers for Large Output Spaces

NeurIPS 2019poster

In extreme classification settings, embedding-based neural network models are currently not competitive with sparse linear and tree-based methods in terms of accuracy. Most prior works attribute this poor performance to the low-dimensional bottleneck in embedding-based methods. In this paper, we dem…

Cited by 75SourcePDFScholar
2019

Dual Variational Generation for Low Shot Heterogeneous Face Recognition

NeurIPS 2019spotlight

Heterogeneous Face Recognition (HFR) is a challenging issue because of the large domain discrepancy and a lack of heterogeneous data. This paper considers HFR as a dual generation problem, and proposes a novel Dual Variational Generation (DVG) framework. It generates large-scale new paired heterogen…

2019

M2FPA: A Multi-Yaw Multi-Pitch High-Quality Dataset and Benchmark for Facial Pose Analysis

ICCV 2019poster

Facial images in surveillance or mobile scenarios often have large view-point variations in terms of pitch and yaw angles. These jointly occurred angle variations make face recognition challenging. Current public face databases mainly consider the case of yaw variations. In this paper, a new large-s…

Cited by 45PDFcodeScholar
2017

Multiscale Quantization for Fast Similarity Search

NeurIPS 2017poster

We propose a multiscale quantization approach for fast similarity search on large, high-dimensional datasets. The key insight of the approach is that quantization methods, in particular product quantization, perform poorly when there is large variance in the norms of the data points. This is a commo…

Cited by 85SourcePDFScholar
2016

Spatial feature learning for robust binaural sound source localization using a composite feature vector

ICASSP 2016accepted

The performance of binaural speech source localization systems can be significantly impacted by an imperfect selection of spatial localization cues, due to the limited bandwidth of speech, and the effects of noise. In order to mitigate these impacts, this paper presents a novel method that combines…

Cited by 0SourceScholar
2015

Binaural localization of speech sources in 3-D using a composite feature vector of the HRTF

ICASSP 2015accepted

Binaural localization of speech sources in 3-D, using head-related transfer functions (HRTFs), always suffers elevation ambiguity due to the limited high frequency spectral information available at the receivers. This paper presents a method that overcomes this limitation by exploiting the interaura…

Cited by 0SourceScholar