← Search

Xiang Xu

20 accepted papers

2026

Aligning Collaborative View Recovery and Tensorial Subspace Learning via Latent Representation for Incomplete Multi-View Clustering

ICLR 2026poster

Multi-view data usually suffer from partially missing views in open scenarios, which inevitably degrades clustering performance. The incomplete multi-view clustering (IMVC) has attracted increasing attention and achieved significant success. Although existing imputation-based IMVC methods perform we…

Cited by 0SourceScholar
2026

Decoupling Vision and Language: Codebook Anchored Visual Adaptation

CVPR 2026

Large Vision-Language Models (LVLMs) use their vision encoders to translate images into representations for downstream reasoning, but the encoders often underperform in domain-specific visual tasks such as medical image diagnosis or fine-grained classification, where representation errors can cascad

Cited by 0SourceScholar
2026

U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences

CVPR 2026

Modeling dynamic 3D environments from LiDAR sequences is central to building reliable 4D worlds for autonomous driving and embodied AI. Existing generative frameworks, however, often treat all spatial regions uniformly, overlooking the varying uncertainty across real-world scenes. This uniform gener

Cited by 0SourcecodeScholar
2026

Veila: Panoramic LiDAR Generation from a Monocular RGB Image

ICRA 2026poster

Realistic and controllable panoramic LiDAR data generation is critical for scalable 3D perception in autonomous driving and robotics. Existing methods either perform unconditional generation with poor controllability or adopt text-guided synthesis, which lacks fine-grained spatial control. Leveragin…

2025

Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR Representations

ICCV 2025poster

LiDAR representation learning aims to extract rich structural and semantic information from large-scale, readily available datasets, reducing reliance on costly human annotations. However, existing LiDAR representation strategies often overlook the inherent spatiotemporal cues in LiDAR sequences, li…

2025

EventFly: Event Camera Perception from Ground to the Sky

CVPR 2025poster

Cross-platform adaptation in event-based dense perception is crucial for deploying event cameras across diverse settings, such as vehicles, drones, and quadrupeds, each with unique motion dynamics, viewpoints, and class distributions. In this work, we introduce EventFly, a framework for robust cross…

Cited by 0SourcePDFScholar
2025

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels

CVPR 2025poster

This work presents a simple yet effective workflow for automatically scaling instruction-following data to elicit pixel-level grounding capabilities of VLMs under complex instructions. In particular, we address five critical real-world challenges in text-instruction-based grounding: hallucinated ref…

Cited by 0SourcePDFScholar
2025

LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes

CVPR 2025poster

LiDAR data pretraining offers a promising approach to leveraging large-scale, readily available datasets for enhanced data utilization. However, existing methods predominantly focus on sparse voxel representation, overlooking the complementary attributes provided by other LiDAR representations. In t…

2025

Model Diagnosis and Correction via Linguistic and Implicit Attribute Editing

CVPR 2025poster

How can we troubleshoot a deep visual model, i.e., understand why it makes certain mistakes and further take action to correct its behavior? We design a Model Diagnosis and Correction system (MDC), an automated framework that analyzes the pattern of errors, proposes candidate causes of attributes, c…

Cited by 0SourcePDFScholar
2025

Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing

CVPR 2025poster

Developing a face anti-spoofing model that meets the security requirements of clients worldwide is challenging due to the domain gap between training datasets and the diverse end-user test data. Moreover, for security and privacy reasons, it is undesirable for clients to share a large amount of thei…

Cited by 0SourcePDFScholar
2025

Salient Concept-Aware Generative Data Augmentation

NeurIPS 2025poster

Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts. This challenge arises because representations in the synthesis…

Cited by 0SourceScholar
2024

FHT-Map: Feature-Based Hybrid Topological Map for Relocalization and Path Planning

RA-L 2024

Topological maps are favorable for their small storage compared to geometric maps. However, they are limited in relocalization and path planning capabilities. To solve the problem, a feature-based hybrid topological map (FHT-Map) is proposed along with a real-time map construction algorithm based on

Cited by 8SourcecodeScholar
2024

RHAML: Rendezvous-Based Hierarchical Architecture for Mutual Localization

RA-L 2024

Mutual localization serves as the foundation for collaborative perception in multi-robot systems. Effectively utilizing limited onboard sensors for mutual localization between marker-less robots is worthwhile. However, due to inadequate consideration of large scale variations of the robot and locali

Cited by 2SourceScholar
2023

Hierarchical Neural Coding for Controllable CAD Model Generation

ICML 2023poster

This paper presents a novel generative model for Computer Aided Design (CAD) that 1) represents high-level design concepts of a CAD model as a three-level hierarchical tree of neural codes, from global part arrangement down to local curve geometry; and 2) controls the generation or completion of CAD…

2022

SkexGen: Autoregressive Generation of CAD Construction Sequences with Disentangled Codebooks

ICML 2022spotlight

We present SkexGen, a novel autoregressive generative model for computer-aided design (CAD) construction sequences containing sketch-and-extrude modeling operations. Our model utilizes distinct Transformer architectures to encode topological, geometric, and extrusion variations of construction seque…

2021

Structured Outdoor Architecture Reconstruction by Exploration and Classification

ICCV 2021poster

This paper presents an explore-and-classify framework for structured architectural reconstruction from aerial image. Starting from a potentially imperfect building reconstruction by an existing algorithm, our approach 1) explores the space of building models by modifying the reconstruction via heuri…

Cited by 15PDFcodeScholar
2019

Adversarial Representation Learning for Text-to-Image Matching

ICCV 2019poster

For many computer vision applications such as image captioning, visual question answering, and person search, learning discriminative feature representations at both image and text level is an essential yet challenging problem. Its challenges originate from the large word variance in the text domain…

Cited by 267PDFcodeScholar
2019

d-SNE: Domain Adaptation Using Stochastic Neighborhood Embedding

CVPR 2019oral

On the one hand, deep neural networks are effective in learning large datasets. On the other, they are inefficient with their data usage. They often require copious amount of labeled-data to train their scads of parameters. Training larger and deeper networks is hard without appropriate regularizati…

Cited by 163PDFcodeScholar
2018

Deep Imbalanced Attribute Classification using Visual Attention Aggregation

ECCV 2018poster

For many computer vision applications, such as image description and human identification recognizing the visual attributes of humans is an essential yet challenging problem. Its challenges originate from its multi-label nature, the large underlying class imbalance and the lack of spatial annotation…

Cited by 274SourcePDFScholar