← Search

Xu Zhou

14 accepted papers

2026

Adaptive Agent Selection and Interaction Network for Image-to-Point Cloud Registration

AAAI 2026technical

Typical detection-free methods for image-to-point cloud registration leverage transformer-based architectures to aggregate cross-modal features and establish correspondences. However, they often struggle under challenging conditions, where noise disrupts similarity computation and leads to incorrect

Cited by 0SourcePDFScholar
2026

Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object Tracking

CVPR 2026

3D single object tracking (SOT) in point clouds is essential for real-world 3D perception, yet it remains challenging due to data sparsity and large variations in scale and structure across diverse object categories. Most existing methods rely on a category-specific paradigm that trains separate mod

Cited by 0SourceScholar
2026

Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning

CVPR 2026

In AI-generated image detection, current cutting-edge methods typically adapt pre-trained foundation models through partial-parameter fine-tuning. However, these approaches often struggle to generalize to forgeries from unseen generators, as the fine-tuned models capture only limited patterns from t

Cited by 0SourceScholar
2025

Generative Map Priors for Collaborative BEV Semantic Segmentation

CVPR 2025poster

Collaborative perception aims to address the constraint of single-agent perception by exchanging information among multiple agents. Previous works primarily focus on collaborative object detection, exploring compressed transmission and fusion prediction tailored to sparse object features. However, t…

Cited by 0SourcePDFScholar
2025

Implicit Correspondence Learning for Image-to-Point Cloud Registration

CVPR 2025highlight

Image-to-point cloud registration aims to estimate the camera pose of a given image within a 3D scene point cloud. In this area, matching-based methods have achieved leading performance by first detecting the overlapping region, then matching point and pixel features learned by neural networks and f…

Cited by 0SourcePDFScholar
2025

InstructSAM: A Training-free Framework for Instruction-Oriented Remote Sensing Object Recognition

NeurIPS 2025poster

Language-guided object recognition in remote sensing imagery is crucial for large-scale mapping and automated data annotation. However, existing open-vocabulary and visual grounding methods rely on explicit category cues, limiting their ability to handle complex or implicit queries that require adva…

Cited by 0SourcecodeScholar
2025

Point Cluster: A Compact Message Unit for Communication-Efficient Collaborative Perception

ICLR 2025poster

The objective of the collaborative perception task is to enhance the individual agent's perception capability through message communication among neighboring agents. A central challenge lies in optimizing the inherent trade-off between perception ability and communication cost. To tackle this bottle…

Cited by 0SourcePDFScholar
2025

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

CVPR 2025poster

Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centric Transformer architecture, where object queries are derived from audio features. However, audio-centric Transformers su…

Cited by 0SourcePDFScholar
2025

TOTF: Missing-Aware Encoders for Clustering on Multi-View Incomplete Attributed Graphs

IJCAI 2025

As the network data in real life become multi-modal and multi-relational, multi-view attributed graphs have garnered significant attention. Numerous methods have achieved excellent performance in multi-view attributed graph clustering; however, they cannot efficiently handle incomplete attribute sce

Cited by 0SourcePDFScholar
2025

Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation

CVPR 2025poster

To tackle the threat of fake news, the task of detecting and grounding multi-modal media manipulation (DGM4) has received increasing attention. However, most state-of-the-art methods fail to explore the fine-grained consistency within local content, usually resulting in an inadequate perception of d…

2025

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

CVPR 2025poster

The advancement of Large Vision Language Models (LVLMs) has significantly improved multimodal understanding, yet challenges remain in video reasoning tasks due to the scarcity of high-quality, large-scale datasets. Existing video question-answering (VideoQA) datasets often rely on costly manual anno…

2024

Aggregation and Purification: Dual Enhancement Network for Point Cloud Few-shot Segmentation

IJCAI 2024poster

Point cloud few-shot semantic segmentation (PC-FSS) aims to segment objects within query samples of new categories given only a handful of annotated support samples. Although PC-FSS demonstrates enhanced category generalization capabilities compared to the fully supervised paradigm, the prevalent…

Cited by 6SourcePDFScholar
2024

DN-4DGS: Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene Rendering

NeurIPS 2024poster

Dynamic scenes rendering is an intriguing yet challenging problem. Although current methods based on NeRF have achieved satisfactory performance, they still can not reach real-time levels. Recently, 3D Gaussian Splatting (3DGS) has garnered researchers' attention due to their outstanding rendering q…