← Search

Xin Wei

17 accepted papers

2026

PSMix: Robust Point Cloud Recognition through Spectral Domain Mixing

ICML 2026poster

While data augmentation is essential for robust point cloud recognition, conventional spatial mixup strategies often compromise geometric integrity by generating physically unrealistic samples. To overcome this limitation, we propose PSMix, which shifts the mixing paradigm to the spectral domain via…

Cited by 0SourceScholar
2026

Robust Cross-Modal Retrieval via Generative Semantic Refinement and Exclusion-Guided Adaptation

ICML 2026poster

Vision-Language Pre-trained (VLP) models are vulnerable to real-world query noise. Current cross-modal Test-Time Adaptation (TTA) methods often rely on high-confidence predictions, which induces confirmation bias and neglects the informative signals in ambiguous Low-Confidence Queries. To address th…

Cited by 0SourceScholar
2026

When Lines Meet Textures: Spatial-Frequency Aligned Diffusion Features for Cross-Sparsity Correspondence

CVPR 2026

Establishing accurate correspondence between sparse line representations and rich textured imagery remains a formidable challenge. While diffusion features excel in semantic correspondence, they struggle to bridge the fundamental gap between abstract sketches and texture-rich photographs. We identif

Cited by 0SourcecodeScholar
2025

3D Test-time Adaptation via Graph Spectral Driven Point Shift

ICCV 2025poster

While test-time adaptation (TTA) methods effectively address domain shifts by dynamically adapting pre-trained models to target domain data during online inference, their application to 3D point clouds is hindered by their irregular and unordered structure. Current 3D TTA methods often rely on compu…

Cited by 0SourcePDFScholar
2025

TriDE-Net: Triple-Densely Extraction Network for Precise Skin Lesion Segmentation

ICASSP 2025accepted

Accurate skin lesion segmentation is crucial for the quantitative analysis of skin cancer. Despite the significant advancements achieved by the deep-learning methods, the segmentation of skin lesions with irregular shapes and significant size variations is still challenging. To address the problem,…

Cited by 0SourceScholar
2025

UMotion: Uncertainty-driven Human Motion Estimation from Inertial and Ultra-wideband Units

CVPR 2025highlight

Sparse wearable inertial measurement units (IMUs) have gained popularity for estimating 3D human motion. However, challenges such as pose ambiguity, data drift, and limited adaptability to diverse bodies persist. To address these issues, we propose UMotion, an uncertainty-driven, online fusing-all s…

2023

Feature Distribution Fitting with Direction-Driven Weighting for Few-Shot Images Classification

AAAI 2023technical

Few-shot learning has received increasing attention and witnessed significant advances in recent years. However, most of the few-shot learning methods focus on the optimization of training process, and the learning of metric and sample generating networks. They ignore the importance of learning the…

Cited by 7SourcePDFScholar
2022

GCLO: Ground Constrained LiDAR Odometry with Low-drifts for GPS-denied Indoor Environments

ICRA 2022poster

LiDAR is widely adopted in Simultaneous Localization And Mapping (SLAM) and High Definition (HD) map production. The accuracy of LiDAR Odometry (LO) is of great importance, especially in GPS-denied environments. However, we found typical LO results are prone to drift upwards along the vertical direc…

Cited by 33SourceScholar
2022

Improving Zero-Shot Entity Linking Candidate Generation with Ultra-Fine Entity Type Information

COLING 2022main

Entity linking, which aims at aligning ambiguous entity mentions to their referent entities in a knowledge base, plays a key role in multiple natural language processing tasks. Recently, zero-shot entity linking task has become a research hotspot, which links mentions to unseen entities to challenge…

2022

Learning Generalizable Part-based Feature Representation for 3D Point Clouds

NeurIPS 2022accept

Deep networks on 3D point clouds have achieved remarkable success in 3D classification, while they are vulnerable to geometry variations caused by inconsistent data acquisition procedures. This results in a challenging 3D domain generalization (3DDG) problem, that is to generalize a model trained on…

2022

Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension

ACL 2022long

Procedural Multimodal Documents (PMDs) organize textual instructions and corresponding images step by step. Comprehending PMDs and inducing their representations for the downstream reasoning tasks is designated as Procedural MultiModal Machine Comprehension (M3C). In this study, we approach Procedur…

2021

GMH: A General Multi-hop Reasoning Model for KG Completion

EMNLP 2021main

Knowledge graphs are essential for numerous downstream natural language processing applications, but are typically incomplete with many facts missing. This results in research efforts on multi-hop reasoning task, which can be formulated as a search process and current models typically perform short…

Cited by 17SourcePDFScholar
2021

Learning Canonical View Representation for 3D Shape Recognition With Arbitrary Views

ICCV 2021poster

In this paper, we focus on recognizing 3D shapes from arbitrary views, i.e., arbitrary numbers and positions of viewpoints. It is a challenging and realistic setting for view-based 3D shape recognition. We propose a canonical view representation to tackle this challenge. We first transform the origi…

Cited by 22PDFcodeScholar
2020

Deep Positional and Relational Feature Learning for Rotation-Invariant Point Cloud Analysis

ECCV 2020poster

In this paper we propose a rotation-invariant deep network for point clouds analysis. Point-based deep networks are commonly designed to recognize roughly aligned 3D shapes based on point coordinates, but suffer from performance drops with shape rotations. Some geometric features, e.g., distances an…

Cited by 44SourcePDFScholar
2020

Object-based Illumination Estimation with Rendering-aware Neural Networks

ECCV 2020poster

We present a scheme for fast environment light estimation from the RGBD appearance of individual objects and their local image areas. Conventional inverse rendering is too computationally demanding for real-time applications, and the performance of purely learning-based techniques may be limited by…

Cited by 29SourcePDFScholar