← Search

Yifan Xie

11 accepted papers

2026

FlexiCup: Wireless Multimodal Suction Cup With Dual-Zone Vision-Tactile Sensing

RA-L 2026

Conventional suction cups lack sensing capabilities for contact-aware manipulation in unstructured environments. This paper presents FlexiCup, a multimodal suction cup with wireless electronics that integrate dual-zone vision-tactile sensing. The central zone dynamically switches between vision and

Cited by 0SourceScholar
2026

ViSA-Gait: Leveraging Vision Foundation Models for Semantic Anchored Gait Recognition

IJCAI 2026

Gait recognition has achieved remarkable success in constrained environments, yet its performance often degrades significantly in cross-domain and cross-vertical-view scenarios. This is primarily due to the fact that domain-specific silhouette geometry causes models to overfit to extrinsic geometric

Cited by 0Scholar
2025

GaussianPU: Color Point Cloud Upsampling via 3D Gaussian Splatting

IROS 2025

Dense colored point clouds enhance visual perception and are of significant value in various robotic applications. However, existing learning-based point cloud upsampling methods are constrained by computational resources and batch processing strategies, which often require subdividing point clouds

Cited by 1SourceScholar
2025

Multiple Rotation Averaging with Constrained Reweighting Deep Matrix Factorization

ICRA 2025

Multiple rotation averaging plays a crucial role in computer vision and robotics domains. The conventional optimization-based methods optimize a nonlinear cost function based on certain noise assumptions, while most previous learning-based methods require ground truth labels in the supervised traini

Cited by 0SourceScholar
2025

Observation-Graph Interaction and Key-Detail Guidance for Vision and Language Navigation

IROS 2025

Vision and Language Navigation (VLN) requires an agent to navigate through environments following natural language instructions. However, existing methods often struggle with effectively integrating visual observations and instruction details during navigation, leading to suboptimal path planning an

Cited by 2SourceScholar
2025

PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis

AAAI 2025technical

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and identity-consistent talking heads from just a few minutes of tr…

Cited by 4SourcePDFScholar
2025

Universal Visuo-Tactile Video Understanding for Embodied Interaction

NeurIPS 2025poster

Tactile perception is essential for embodied agents to understand the physical attributes of objects that cannot be determined through visual inspection alone. While existing methods have made progress in visual and language modalities for physical understanding, they fail to effectively incorporate…

Cited by 0SourceScholar
2024

Cross-Modal Information-Guided Network Using Contrastive Learning for Point Cloud Registration

RA-L 2024

The majority of point cloud registration methods currently rely on extracting features from points. However, these methods are limited by their dependence on information obtained from a single modality of points, which can result in deficiencies such as inadequate perception of global features and a

Cited by 14SourcecodeScholar
2024

Hyperbolic Image-and-Pointcloud Contrastive Learning for 3D Classification

IROS 2024poster

3D contrastive representation learning has exhibited remarkable efficacy across various downstream tasks. However, existing contrastive learning paradigms based on cosine similarity fail to deeply explore the potential intra-modal hierarchical and cross-modal semantic correlations about multi-modal…

Cited by 0SourceScholar
2024

Matching Distance and Geometric Distribution Aided Learning Multiview Point Cloud Registration

RA-L 2024

Multiview point cloud registration plays a crucial role in robotics, automation, and computer vision fields. This letter concentrates on pose graph construction and motion synchronization within multiview registration. Previous methods for pose graph construction often pruned fully connected graphs

Cited by 8SourcecodeScholar