← Search

Xiaotian Yin

7 accepted papers

2026

Adaptive Agent Selection and Interaction Network for Image-to-Point Cloud Registration

AAAI 2026technical

Typical detection-free methods for image-to-point cloud registration leverage transformer-based architectures to aggregate cross-modal features and establish correspondences. However, they often struggle under challenging conditions, where noise disrupts similarity computation and leads to incorrect

Cited by 0SourcePDFScholar
2026

Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models

ICLR 2026poster

Vision-Language Models (VLMs), such as CLIP, have shown strong zero-shot generalization but remain highly vulnerable to adversarial perturbations, posing serious risks in real-world applications. Test-time defenses for VLMs have recently emerged as a promising and efficient approach to defend agains…

Cited by 0SourceScholar
2026

FS-I2P: A Hierarchical Focus–Sweep Registration Network with Dynamically Allocated Depth

ICML 2026poster

Image-to-point cloud registration is often challenged by viewpoint changes, cross-modal discrepancies, and repetitive textures, which induce scale ambiguity and consequently lead to erroneous correspondences. Recent detection-free methods alleviate this issue by leveraging multi-scale features and t…

Cited by 0SourceScholar
2026

Rethinking 2D-3D Registration: A Novel Network for High-Value Zone Selection and Representation Consistency Alignment

CVPR 2026

Both detection-then-match and detection-free methods have been extensively studied for image-to-point cloud registration, yet they still face significant challenges. The detection-then-match approach emphasizes high-quality correspondences but is limited by the availability of repeatable keypoints,

Cited by 0SourceScholar
2025

CA-I2P: Channel-Adaptive Registration Network with Global Optimal Selection

ICCV 2025poster

Detection-free methods typically follow a coarse-to-fine pipeline, extracting image and point cloud features for patch-level matching and refining dense pixel-to-point correspondences. However, differences in feature channel attention between images and point clouds may lead to degraded matching res…

Cited by 0SourcePDFScholar
2025

Exploring the Better Multimodal Synergy Strategy for Vision-Language Models

AAAI 2025technical

Vision-Language models (VLMs) have shown great potential in enhancing open-world visual concept comprehension. Recent researches focus on an optimum multimodal collaboration strategy that significantly advances CLIP-based few-shot tasks. However, existing prompt-based solutions suffer from unidirect…

Cited by 0SourcePDFScholar
2024

Task-Adaptive Prompted Transformer for Cross-Domain Few-Shot Learning

AAAI 2024technical

Cross-Domain Few-Shot Learning (CD-FSL) aims at recognizing samples in novel classes from unseen domains that are vastly different from training classes, with few labeled samples. However, the large domain gap between training and novel classes makes previous FSL methods perform poorly. To address t…